DeepSeek V4 Models, Harness, and API Discount Windows: The Complete Guide (2026)

TL;DR

DeepSeek had a very big month — and then another one. It open-sourced a small model that embarrasses its own flagship, shipped an open-source Claude Code rival called DeepSeek Harness, and started billing its API like an electricity company — half price for most hours. On September 10, 2026 it replaced the Flash tier with DeepSeek-V4.1-Flash (deepseek-flash): a 552B-parameter model with native image understanding that DeepSeek says beats its own V4-Pro on performance, cost and speed, at prices cut 9–57%. This guide explains the current API lineup, gets Harness running on a Mac in four commands, and gives you every discount window and price — all dated and sourced. The bigger picture of how China came to set the world's price floor got its own article: Why China Is Winning the AI Race.

1.6T
Pro total params
49B active / token
552B
V4.1-Flash params
8B/16B active, native vision
1M
context window
384K max output
50%
off-peak discount
weekdays + all weekend
$0.60
Flash output off-peak /1M
was $0.66 before Sep 10
MIT
license (both tiers)
open weights on HF

DeepSeek-V4.1-Flash: the model that replaced the flagship-killer (and now has eyes)

Update — September 10, 2026: V4-Flash is retired, V4.1-Flash is the Flash tier

On September 10, 2026 at 04:00 UTC, DeepSeek released DeepSeek-V4.1-Flash, the first model in a new architecture family — and retired V4-Flash-0731 and V4-Flash-Vision-Exp on the API. The new canonical model name is deepseek-flash; the old names (deepseek-v4-flash, deepseek-v4-flash-vision-exp) still work for compatibility but serve V4.1-Flash and bill at Flash prices. DeepSeek says "tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime" — and cut Flash prices 9–57% at the same time. Everything below the spec sheet describes the retired V4-Flash-0731 and is kept as historical context.

Every so often a "small" model comes along that makes the flagship look overpriced. On July 31, 2026, DeepSeek released the production version of V4-Flash and posted the weights on Hugging Face under the MIT license — they promptly hit the top of the trending charts. The 0731 in the name is just the release date: July 31. Independent testers found it beating DeepSeek's own flagship V4-Pro on several agentic benchmarks — at a fraction of the cost.

On September 10, 2026, DeepSeek went further and replaced the Flash tier outright with V4.1-Flash: a 552B-parameter causal encoder-decoder MoE (8B active on prefill, 16B on decode) with native image understanding — the first multimodal model DeepSeek has shipped on the API — a 1M-token context, 384K max output, and roughly a quarter of the previous KV-cache footprint. The weights are MIT-licensed on Hugging Face. The old Flash's fine-tune story (below) explains how a small model can outrun a flagship; V4.1-Flash is the same thesis with a new architecture and eyes.

The V4.1-Flash spec sheet (current)

V4.1-Flash at a glance
SpecificationDeepSeek-V4.1-Flash (deepseek-flash)
Model typeSparse Mixture-of-Experts (MoE), causal encoder-decoder (20+20 layers, 384 routed experts + 1 shared, 6 active)
Total parameters552B
Active parameters per token8B (prefill) / 16B (decode)
Context window1,000,000 tokens (1M)
Max output tokens384K
Input modalitiesText + native image understanding (multimodal); images are converted to tokens and billed with text input
LicenseMIT — weights free for commercial and non-commercial use
API model namedeepseek-flash — legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp still resolve but route here
Concurrency limit2,500 concurrent requests
Efficiency~1/4 the KV-cache footprint of the previous V4-Flash, DSpark speculative decoding built in

A million tokens of context, MIT license, native vision, and prices that undercut most of the market. Only 8B–16B of the 552B parameters wake up for any given token — like a 552,000-person company where a few thousand people touch any one project. That's what makes it cheap to run.

The retired V4-Flash-0731 spec sheet (historical)

V4-Flash-0731 at a glance (historical)
SpecificationDeepSeek-V4-Flash-0731 (retired Sep 10)
Model typeSparse Mixture-of-Experts (MoE) transformer
Total parameters284B (304B with the attached DSpark speculative-decoding draft module)
Active parameters per token13B
Context window1,000,000 tokens (1M)
Max output tokens384K (~122.7 tokens/sec generation)
LicenseMIT — weights free for commercial and non-commercial use
ModesThinking (low / high / max reasoning effort) + non-thinking
Tool calls / context cachingBoth supported
Concurrency limit2,500 concurrent requests (high-throughput friendly)
Self-hosting3-bit quantized version runs in ~110 GB of memory; full FP8 weights ≈ 166.9 GB

How it handles a million tokens without the bill exploding

The expensive part of long-context AI is the key-value (KV) cache — attention's working memory, which grows with every token you feed in. V4 attacks it from two directions. Some attention layers keep detailed notes, condensing every 4 tokens into one entry and only re-reading the relevant ones (CSA, Compressed Sparse Attention). Other layers keep summaries, condensing every 128 tokens into one entry that always gets read (HCA, Heavily Compressed Attention).

The payoff, per Goldman Sachs research: at a full 1M-token input, V4-Flash needs only ~10% of the compute and ~7% of the KV-cache memory of DeepSeek V3.2. A few more tricks stack on top:

  • Speculative decoding via the attached DSpark draft module — a small "intern" model drafts several tokens ahead, and the main model checks them all in one pass instead of one at a time. vLLM and SGLang enable it with a single flag: --speculative-config.
  • FP4 quantization for the MoE expert weights and FP8 for everything else, cutting KV-cache memory by up to 90%.
  • A packaging change: Jinja chat templates are gone, replaced by a Python script package (encoding_dsv4) and a configurable reasoning_effort parameter (low / high / max).

How it was trained (the interesting bit)

Flash-0731 was pretrained on more than 32 trillion tokens, then fine-tuned in two stages. First, DeepSeek built a separate specialist model per domain — math, coding, agentic tasks — using supervised fine-tuning plus GRPO reinforcement learning. Then it merged the 10+ specialists via on-policy distillation: the merged model wrote its own answers, and training nudged each one toward how the best specialist would have answered. That's how one model inherits ten skills without averaging them into mush.

Two details worth knowing. The three reasoning levels (low / high / max) were trained as genuinely distinct behaviors with different length penalties and context windows — "max" prepends a system prompt pushing the model to fully decompose problems and test edge cases. And during tool-using agentic tasks, Flash keeps its entire reasoning history in context across every round, including across user messages — something V3.2 simply threw away.

What independent benchmarks say (August 2026)

Independent results, max reasoning effort
BenchmarkV4-Flash-0731 (max reasoning)Context
Artificial Analysis Intelligence Index50 (preview: 40, V4-Pro: 44)Ties Gemini 3.6 Flash (50); 1 pt behind GPT-5.6 Luna & GLM-5.2 (51); leader is Kimi K3 (57)
GDPval-AA v2 (real-work head-to-head)1,558 Elo2nd-best open-weights model (behind Kimi K3's 1,685; ahead of GLM-5.2's 1,508)
Terminal-Bench 2.1 (CLI agent tasks)82.7%+21 pts vs. April preview (61.8%)
τ³-Bench Banking (multi-turn tool use)31.1%~8 pts above preview
CodeArena WebDev (front-end dev)1,5777th overall, 3rd among open-weights models
Cost per Intelligence-Index task$0.03vs. $0.05 for the similar-intelligence GPT-5.6 Luna

The story in one line, per DeepLearning.AI's The Batch: Flash-0731 sits on Artificial Analysis' Pareto frontier for intelligence vs. cost per task — no tracked model is both smarter and cheaper. It also landed during a brutal pricing week: OpenAI cut GPT-5.6 Luna by 80%, Google shipped Gemini 3.6 Flash, and Thinking Machines released Inkling Small. The market's center of gravity is now intelligence-per-dollar.

Cost per Intelligence-Index task
GPT-5.6 Luna
$0.05
DeepSeek V4-Flash Pareto frontier
$0.03

What it cost before the price change

At launch, the first-party API charged a flat $0.0028 per million cached-input tokens, $0.14 per million input tokens, and $0.28 per million output tokens (≈¥0.02 / ¥1 / ¥2). The August 16 tariff overhaul replaced that flat rate with the peak/off-peak schedule — and the September 10 release of V4.1-Flash cut Flash prices again, between 9% (output) and 57% (cache-hit input) below the August card. Current numbers are in the pricing section.


DeepSeek-V4-Pro-0813: the flagship, with caveats

Update — September 14, 2026: Pro stays (the reroute was reversed)

DeepSeek initially announced that all deepseek-v4-pro requests would route to V4.1-Flash at Flash prices from September 14, 04:00 UTC until "V4.1-Pro" launches. That plan was reversed before it took effect: the API docs now state, "in response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." Pro keeps its own rates (below), and there is still no date for V4.1-Pro.

Pro is the 1.6-trillion-parameter flagship, and its August 13, 2026 "GA Release" was already the third V4-Pro drop in four months: an open preview on April 24, GA on July 19, then the quieter 0813 production build on August 13. Per DeepSeek's own change log, 0813 is mostly a serving-efficiency update — ~51.7B new parameters across four new DSpark speculative-decoding keys, not a new brain. On the API it simply serves as deepseek-v4-pro; no model-name migration needed.

The spec sheet

Pro-0813 at a glance
SpecificationDeepSeek-V4-Pro-0813
Model typeSparse Mixture-of-Experts (MoE)
Total parameters1.6 trillion
Active parameters per token49 billion
Context window1,000,000 tokens (1M)
Max output tokens384K
LicenseMIT, open weights on Hugging Face (full FP8 weights ≈ 892.7 GB — not single-workstation self-hostable)
ModesThinking (default) + non-thinking; reasoning effort low / high / max
API compatibilityOpenAI format (https://api.deepseek.com), Responses API, Anthropic API (/anthropic), JSON output, tool calls, Chat Prefix Completion (Beta), FIM Completion (Beta, non-thinking only)
Concurrency limit500 concurrent requests
Headline use caseAgentic coding; native Codex integration with one-click setup; "Expert Mode" in the DeepSeek app/web

What's actually new under the hood

Two ideas carry the 1.6T model. mHC (Manifold-Constrained Hyper-Connections) keeps signal amplification under 2×, which keeps training a 1.6T model stable at only ~6.7% compute overhead. DSA (DeepSeek Sparse Attention) is the same hybrid compressed/hierarchical attention as Flash — CSA plus HCA — pushing context-cost growth closer to linear than quadratic. The result, per Goldman Sachs and DeepSeek's own figures: at the 1M-token setting, V4-Pro needs ~27% of the single-token inference compute and ~10% of the KV cache that V3.2 needed. The 0813 build then spends its ~51.7B new DSpark parameters on speculative decoding, i.e., faster and cheaper serving.

The benchmark reality check

DeepSeek's own numbers — run through DeepSeek Harness at max reasoning effort, versus the April preview — are big:

Vendor-reported deltas vs. April preview
BenchmarkV4-Pro-0813 (vendor)Δ vs. preview
Terminal-Bench 2.187.9+15.8
DeepSWE62.7+49.9
CyberGym83.3+30.6
Toolathlon-Verified74.1+18.2
NL2Repo61.5+23.0
SWE-bench Verified>80% (official model card)—

Independent labs (AIToolsReview, Aug 18, 2026) tell a more nuanced story:

Third-party measurements
Independent testResultTakeaway
Artificial Analysis Intelligence Index53 — #3 of ~106 tracked models (median 27)Genuinely frontier-ranked on composite intelligence
CoderSera SWE-bench Verified (neutral harness)96.40% ±0.83 — #2 (behind Claude Opus 5's 97.00%)Excellent real coding score at a fraction of the cost
CoderSera Terminal-Bench 2.1 (neutral harness)54.68% vs. vendor's 87.933-point vendor/neutral gap — and V4-Flash scored 67.04% on the same harness, beating Pro
Artificial Analysis AA-Omniscience (honesty)0.83 (near floor; Claude Opus 5: 37.07)The model almost never says "I don't know" — weak calibration
NIST CAISI (April preview build)94% jailbreak compliance vs. 8% for US reference models; ~8-month capability gap vs. frontierMost recent independent safety data; not yet re-run on 0813
Terminal-Bench 2.1 — vendor vs. neutral harness
V4-Pro-0813 (vendor) DeepSeek-reported
87.9%
V4-Flash-0731 (neutral) CoderSera — beats Pro
67.04%
V4-Pro-0813 (neutral) CoderSera
54.68%

The Terminal-Bench gap is the one to sit with: on a neutral harness, the smaller Flash scored 67.04% while Pro managed 54.68%. Treat vendor magnitudes as company-reported — the direction (large agentic gains) is credible and consistent across five benchmarks, but run your own evals before committing.

What it's like to actually use

Hands-on testing (via AIToolsReview) scored V4-Pro at 76.25% (61/80) on an 8-question practical test, with full marks on a hard math problem and a long-horizon agentic task. Front-end and one-shot UI generation are its clearest strength — Three.js scenes, app clones, physics demos. Known weaknesses: it overthinks and overengineers simple tasks, its honesty calibration is weak (see the AA-Omniscience row above), it's text-only, and there's no published agentic-safety evaluation.

Bottom line on 0813

Expect large agentic gains over the preview (credible across five vendor benchmarks), expect V4-Flash to beat Pro on many simple tasks, and add a verification step for factual work. That's the honest read.


The wider war

Why is DeepSeek able to set the price floor the whole market now follows? That story — China's open-weights strategy, structural cost advantage, and domestic-silicon hedge — got its own article: Why China Is Winning the AI Race (2026). The rest of this guide stays focused on the models, the harness, and the prices.


DeepSeek Harness: DeepSeek's answer to Claude Code

On the same day as the Pro GA — August 13, 2026 — DeepSeek quietly shipped something arguably more interesting: DeepSeek Harness (dsh), an MIT-licensed, plugin-first agent framework, on GitHub at deepseek-ai/deepseek-harness. Its README states the organizing idea in three words: "everything is a plugin." DeepSeek used it internally to validate V4-Flash-0731 and to produce V4-Pro's benchmark table. Everyone calls it an open-source Claude Code rival — but the deliberate difference is that it's not locked to DeepSeek models at all.

Status: developer preview

v0.1.0-rc.5 at time of writing. The README warns plainly: there will be compatibility-breaking changes. Not yet a stable production product.

Everything is a plugin — literally

The runtime is built on Cordis, a vendored dependency-injection framework (cordiverse/cordis, designed per A Programming Paradigm for Spatiotemporal Composability). Every part of the product is a plugin — the model adapter, tool registry, session log, agent loop, sandbox, even the web UI — so every part is replaceable from configuration. There's no privileged core to patch: extensions mount beside the other plugins, and registrations are reversible effects that unwind when a plugin unloads.

The docs describe eight independently swappable layers:

eight independently swappable layers — every part replaceable from configuration
01
Inference
Model adapters via ctx.llm — pluggable connection to any configured provider
02
Tools
Tool registry, schemas, policy enforcement, execution pipeline
03
State
Append-only SessionEvent log with persistence, replay, fork, resume
04
Control
Agent registry + loop driver — goals, turns, steps, cancellation
05
Execution
Filesystem, shell, subprocess, terminal, and sandbox providers
06
Composition
Profiles, bundles, patches, and runtime overlays
07
Experience
Bundled Web app, conversation nodes, settings UI
08
Framework
Cordis itself — the DI system underneath everything

It speaks every provider

Despite the branding, the harness supports DeepSeek, Anthropic, OpenAI, Amazon Bedrock, Google Vertex, Azure, and any custom OpenAI-compatible endpoint as interchangeable inference providers. Swapping providers is a configuration change, not a rewrite — which makes dsh infrastructure a team can standardize on regardless of which model wins a given benchmark cycle.

The features that matter

  • Capability seams — each swappable capability has three roles (Service Definition / Provider / Consumer). One provider swap can move Bash, PTY, LSP, and subprocess execution together to a remote sandbox, with no provider forks.
  • The session log is the source of truth — "model-visible ⟺ logged": anything that reaches a model request must be reconstructable from the append-only log. Fork, resume, transcripts, telemetry, and persistence all derive from that stream.
  • Typed event system — durable session events (session/event, turn/*, step/*, tool/*) plus live extension events (agent/*, tools/*) with waterfall semantics, letting plugins intercept, rewrite, or reject requests.
  • Agent loop — steps (one model request + tool calls) compose into turns; agent/pre-step, agent/request, and llm/stream are the interception points.
  • Sandboxing & approvals — filesystem, shell, subprocess, terminal, and sandbox are separate configurable plugins, and an approval layer gates privileged operations under the active permission policy (the Web UI asks before doing anything spicy).
  • Goals, subagents, workflow, plan mode — same-session objectives (ctx.goals), subagent delegation with pluggable providers, workflow orchestration, and plan mode as logged state.
  • Skills & hooks — a skill registry/provider plus Claude Code and Codex hook bridges.
  • Self-modification — the agent can inspect and mount its own plugins (the demo:cordis demo modifies its live runtime).
  • Python SDK — deepseek-harness-sdk drives the bundled runtime as a subprocess over newline-delimited JSON-RPC on stdio.

Ways to run it

  • Web UI — npx @deepseek-ai/dsh web (or pnpm dsh web from source) serves a browser agent interface at http://127.0.0.1:3080.
  • Headless — dsh --profile headless "task" runs one fresh persisted session and prints the final answer.
  • ACP — an automation-only Agent Client Protocol (JSON-RPC stdio) server.
  • Plugin management — dsh plugin --profile <name> <pnpm args> installs out-of-tree plugins per profile.
  • Reported operational modes: standard (general tasks), code-focused (multi-app automation), creative (custom tools), and minimal (isolated testing — reportedly how DeepSeek tested V4-Flash before release).

Installing DeepSeek Harness from source on macOS

Two ways in: npx @deepseek-ai/dsh web if you just want to look around, or from source if you want to hack on it. Here's the from-source route on a Mac.

Prerequisites

macOS prerequisites
RequirementVersion / how to get it
Node.js^22.19.0 or >=24.0.0 (engines field). Install via nodejs.org installer or brew install node@24
pnpm11.7.0 (repo-pinned). With Node installed, enable Corepack: corepack enable (then pnpm --version should resolve; if not, corepack prepare pnpm@11.7.0 --activate)
Git2.26+ (brew install git if needed; Xcode Command Line Tools optional)
DeepSeek API keyOptional for the UI to do real work — get one at platform.deepseek.com and set DEEPSEEK_API_KEY or a repo-root .env
Disk spaceA few GB (workspace deps + built artifacts); V4 model weights are not required (API-driven)
One macOS-specific note

The repo's native/landlock-run workspace (a Landlock self-restrict-then-exec launcher) is Linux-only sandboxing; on macOS it's simply not used, and the harness falls back to the configured non-Landlock providers. No special native toolchain is required for the standard build.

The four commands

install-from-source.sh
# 1. Clone the repository
git clone https://github.com/deepseek-ai/deepseek-harness.git
cd deepseek-harness

# 2. Install dependencies (pnpm workspace; postinstall sets up lefthook git hooks)
pnpm install

# 3. Build everything (host lib -> client lib -> web frontend)
pnpm run build

# 4. Launch the Web UI (serves at http://127.0.0.1:3080 by default)
pnpm dsh web

That's it. Open http://127.0.0.1:3080, head to Settings → Models, and add your DeepSeek API key — the model route becomes usable immediately, no server restart. Then choose a workspace (your project directory), and the session composer unlocks.

Configuration & useful commands

dsh-config.sh
# One-shot headless task from source (needs DEEPSEEK_API_KEY)
pnpm dsh --profile headless "summarize this workspace"

# Inspect the composed plugin tree without booting
pnpm dsh --profile web --dump-config

# Frontend dev with hot reload (HMR) — requires the dev:web watcher running
pnpm run dev:web

# Credentials live in a gitignored .env at the repo root
#   DEEPSEEK_API_KEY=sk-...
#   DEEPSEEK_BASE_URL=https://...   # optional, defaults to public API

When it misbehaves

Common macOS install issues
ProblemFix
pnpm: command not foundcorepack enable after installing Node; or npm install -g pnpm@11.7.0
corepack not presentOlder Node bundles it; otherwise install Node 22.19+ fresh
Build fails on tsdown/tsc OOMIncrease Node heap: export NODE_OPTIONS=--max-old-space-size=8192
Lefthook hook errors after clonenode scripts/install-lefthook.mjs (also needed if postinstall was skipped)
Web UI loads but no model worksAdd API key in Settings → Models; check .env for DEEPSEEK_API_KEY
"Port 3080 in use"pnpm dsh web --port 8080 (app flags come after the launcher's)
You want the packaged binary insteadnpx @deepseek-ai/dsh web — same UI, no build step

DeepSeek API discount windows, off-peak times, and prices

On August 13, 2026, DeepSeek announced something unusual for an LLM API: peak/off-peak (峰谷) tariffing — tokens priced like electricity. It took effect at 16:00 UTC on August 16, 2026 (00:00 Beijing time, August 17), and on August 24 the fine print changed again: peak hours now apply Monday through Friday only, with weekends fully off-peak. The September 10 release of V4.1-Flash then cut Flash prices while keeping the schedule. This is the money section.

The headline answer

The headline

Every hour outside the two weekday peak windows is billed at 50% of the peak price. The peak windows are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday — and all of Saturday and Sunday is off-peak, 24 hours a day.

Peak vs. off-peak — a 24-hour map
Peak (full price) Off-peak (50% off)
UTC
00
01
02
03
04
05
06
07
08
09
10
11
12
13
14
15
16
17
18
19
20
21
22
23
BeijingUTC+8
00
01
02
03
04
05
06
07
08
09
10
11
12
13
14
15
16
17
18
19
20
21
22
23
00:0006:0012:0018:0024:00
  • Peak #1 01:00–04:00 UTC = 09:00–12:00 Beijing
  • Peak #2 06:00–10:00 UTC = 14:00–18:00 Beijing
  • Weekdays only peak applies Mon–Fri — weekends are fully off-peak
  • Off-peak every other weekday hour — 17 hours/day, plus all weekend at 50% of peak

The purple marker shows where you are right now — it moves with live UTC time so you can see at a glance whether the API is billing you full price or half.

The two peak windows (weekdays only)
WindowUTCBeijing (UTC+8)
Peak #1 (full price)01:00–04:0009:00–12:00
Peak #2 (full price)06:00–10:0014:00–18:00
Off-peak (50% of peak)every other weekday hourevery other weekday hour
Weekend (off-peak all day)Sat + Sun, 24hSat + Sun, 24h
  • Total peak time is 7 hours per weekday, 35 hours per week — about 21% of the 168-hour week (it was 49 hours/week before the August 24 change). Everything else — 133 hours a week — is off-peak. Official DeepSeek docs: "Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday (all other hours are off-peak)."
  • Practical read for the US: on Eastern time (summer), peak runs 21:00–00:00 and 02:00–06:00 — your 9-to-5 workday bills off-peak, and the overnight batch job someone scheduled "because nobody is using anything then" is the one that pays double. When US clocks fall back on November 1 (October 25 in the EU), each local window shifts one hour earlier. For China-based teams, lunch (12:00–14:00) and everything after 18:00 Beijing is off-peak.

Current per-1M-token prices (official, checked Sep 28, 2026)

deepseek-flash (V4.1-Flash) — concurrency 2,500. Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp route here and bill at these rates. Image inputs are converted to tokens and billed with text input:

deepseek-flash 2,500 concurrent requests
50% off at off-peak
Token type Off-peak 50% Peak 100%
Input — cache hit $0.003 ≈¥0.02 $0.006 ≈¥0.04
Input — cache miss $0.15 ≈¥1.00 $0.30 ≈¥2.00
Output $0.60 ≈¥4.00 $1.20 ≈¥8.00
per 1M tokens · Off-peak = exactly ½ of peak

deepseek-v4-pro (Pro-0813) — concurrency 500, billing unchanged after the September 14 reroute was reversed:

deepseek-v4-pro 500 concurrent requests
50% off at off-peak
Token type Off-peak 50% Peak 100%
Input — cache hit $0.022 ≈¥0.15 $0.044 ≈¥0.30
Input — cache miss $0.66 ≈¥4.50 $1.32 ≈¥9.00
Output $1.98 ≈¥13.50 $3.96 ≈¥27.00
per 1M tokens · Off-peak = exactly ½ of peak

Off-peak is exactly half of peak in every category. Base URLs: https://api.deepseek.com for OpenAI format, https://api.deepseek.com/anthropic for Anthropic format. Billing deducts from granted balance first, then topped-up balance.

What changed on September 10 (the price cut)

With V4.1-Flash, DeepSeek cut Flash rates between 9% and 57% versus the August card. The cache-hit cut is the steepest — the same stable-prefix workloads got dramatically cheaper:

Flash price cuts, Sep 10, 2026 (same percentages at peak)
Token categoryV4-Flash (Aug 21 rates)V4.1-Flash (Sep 10 rates)Cut
Cache-hit input (off-peak)$0.007$0.003−57%
Cache-miss input (off-peak)$0.22$0.15−32%
Output (off-peak)$0.66$0.60−9%

The honest part: base prices are still above the original flat rate

The original July flat rates for V4-Flash were ¥0.02 / ¥1 / ¥2 (cache-hit-in / cache-miss-in / output) — about $0.0028 / $0.14 / $0.28. The new V4.1-Flash card is much closer to that floor than the August card was, but it hasn't fully returned to it:

deepseek-flash (V4.1-Flash) — increase over the original flat rate:

V4.1-Flash: % vs. original flat rate
Token categoryPeakOff-peak
Cache-hit input+114%+7%
Cache-miss input+114%+7%
Output+329%+114%

deepseek-v4-pro — increase over its GA flat rate (unchanged since August 16):

Pro: % increase vs. old flat rate
Token categoryPeakOff-peak
Cache-hit input+1,100%+500%
Cache-miss input+200%+50%
Output+350%+125%
Read the fine print

"Off-peak" is 50% off the peak price — not a discount versus the original flat rate. For V4.1-Flash, off-peak input is back within ~7% of the original flat rate, but off-peak output still costs ~2.14× it, and peak output ~4.29×. Pro's off-peak output runs ~2.28× its original flat rate, and peak ~4.55×. Plan your cost model on the current table, not the "discount" label.

Three real levers to pull

  1. Shift work to off-peak (50% off — and weekends are all off-peak). Best for batch data processing, bulk translation, data cleaning, offline analysis, overnight model evaluation, report generation — anything with loose latency requirements. Interactive chat, online agents, and real-time copilots generally can't wait. With the August 24 change, a weekend batch pays the off-peak rate all 48 hours.
  2. Maximize cache hits (the biggest multiplier). Cache-hit input is ~50× cheaper than cache-miss for Flash ($0.003 vs. $0.15), ~30× for Pro ($0.022 vs. $0.66). Reuse stable system prompts, long shared contexts, and prefix-stable requests so more of your input tokens land on the cache — and note the September 10 cut hit cache prices hardest (−57%).
  3. Pick the right tier for the job. Flash (deepseek-flash) has 5× the concurrency (2,500 vs. 500), ~3.3× cheaper output, native vision, and — per DeepSeek's own testing — beats Pro on performance, cost, speed and total runtime. Route high-throughput, multimodal, or non-thinking workloads to Flash, and reserve Pro for the complex agentic and front-end work where the 1.6T model's depth still earns its premium.

Why DeepSeek did this

Model inference occupies AI compute around the clock, but traffic concentrates in working hours while idle periods under-utilize the clusters. Pricing tokens like electricity (峰谷电价) uses the price lever to smooth demand, raise utilization, and relieve peak pressure — a sign the LLM price war is entering a "2.0" phase of fine-grained compute-scheduling economics (智东西/Zhidx, Techritual).


Frequently asked questions

QWhat's the difference between deepseek-flash (V4.1-Flash) and deepseek-v4-pro (V4-Pro-0813)?

Since September 10, 2026, the Flash tier is DeepSeek-V4.1-Flash: a 552B-parameter MoE with native image understanding, ~$0.15/$0.60 per million tokens off-peak, 2,500 concurrent requests, and — per DeepSeek's own testing — it beats V4-Pro on performance, cost, speed and total runtime. Pro (1.6T total, 49B active) is the flagship for complex agentic coding and front-end generation, with 500 concurrent requests, text-only input, and unchanged billing after the September 14 reroute was reversed. V4.1-Pro is expected eventually; there is no date.

QWhen are DeepSeek's API discounts available?

Off-peak pricing — 50% of the peak rate — applies Monday through Friday except the two peak windows (01:00–04:00 and 06:00–10:00 UTC, i.e. 09:00–12:00 and 14:00–18:00 Beijing time), and all of Saturday and Sunday is off-peak. That is 7 peak hours per weekday — 35 of 168 hours per week — and 133 off-peak hours.

QIs DeepSeek's off-peak pricing actually cheaper than before?

It depends which "before." Versus the September 10 price cut: off-peak is exactly 50% of the current peak price, and the September cut lowered Flash rates 9–57% versus the August card. Versus the original July flat rate: V4.1-Flash off-peak input is back within ~7% of it, but output still costs ~2.14× off-peak and ~4.29× at peak. Plan on the current table, not the "discount" label.

QWhat happened to deepseek-chat and deepseek-reasoner?

The two legacy model names were fully retired on July 24, 2026 at 15:59 UTC — requests using them now return errors. Migrate to model: "deepseek-flash" (thinking mode is a request parameter, not a separate model name). The V4-era names deepseek-v4-flash and deepseek-v4-flash-vision-exp also route to V4.1-Flash and bill at Flash prices, but pin deepseek-flash going forward.

QIs DeepSeek Harness free, and does it only work with DeepSeek models?

The harness is MIT-licensed and free — you only pay for whichever inference provider you plug in. It supports DeepSeek, Anthropic, OpenAI, Amazon Bedrock, Google Vertex, Azure, and any custom OpenAI-compatible endpoint.

QHow do I install DeepSeek Harness from source on macOS?

git clone https://github.com/deepseek-ai/deepseek-harness.git && cd deepseek-harness && pnpm install && pnpm run build && pnpm dsh web. You will need Node 22.19+/24+, pnpm 11.7.0 (via Corepack), and Git 2.26+. The web UI serves at http://127.0.0.1:3080 by default.

QWhich DeepSeek model is better for coding agents?

DeepSeek says V4.1-Flash beats V4-Pro on performance, cost and speed, and independent tests of the retired V4-Flash already showed Flash beating Pro on a neutral Terminal-Bench harness (67.04% vs. 54.68%). Practical rule: start with deepseek-flash for throughput and cost; keep Pro for the hardest multi-step tasks until V4.1-Pro ships.

QDo DeepSeek models support images, audio, or video input?

Since September 10, 2026, deepseek-flash (V4.1-Flash) supports native image understanding — images are converted to tokens and billed as input. V4-Pro remains text-only.

QAre DeepSeek V4 model weights open source?

Yes. V4.1-Flash and V4-Pro-0813 are both MIT-licensed open weights on Hugging Face. V4.1-Flash is the practical self-host tier; Pro's full FP8 weights are ~892.7 GB — multi-GPU server territory.


Sources & further reading

Official DeepSeek

Models & benchmarks

Harness

Related reading

Pricing & peak/off-peak tariff

Last updated

Last updated: September 28, 2026. Prices and availability verified against DeepSeek's official API documentation and the sources above as of September 28, 2026. Key changes since the previous update: DeepSeek-V4.1-Flash shipped September 10 (new canonical name deepseek-flash, Flash prices cut 9–57%, native vision); peak hours became weekdays-only on August 24 (weekends fully off-peak); the planned September 14 V4-Pro reroute was reversed, so Pro continues at unchanged billing; and the legacy names deepseek-chat/deepseek-reasoner were fully retired July 24. DeepSeek reserves the right to adjust pricing — re-check the official Models & Pricing page before making budget decisions. Vendor-reported benchmark figures should be treated as company-reported until independently reproduced.

← Previous