DeepSeek V4 Models, Harness, and API Discount Windows: The Complete Guide (2026)
DeepSeek had a very big month — and then another one. It open-sourced a small model that embarrasses its own flagship, shipped an open-source Claude Code rival called DeepSeek Harness, and started billing its API like an electricity company — half price for most hours. On September 10, 2026 it replaced the Flash tier with DeepSeek-V4.1-Flash (deepseek-flash): a 552B-parameter model with native image understanding that DeepSeek says beats its own V4-Pro on performance, cost and speed, at prices cut 9–57%. This guide explains the current API lineup, gets Harness running on a Mac in four commands, and gives you every discount window and price — all dated and sourced. The bigger picture of how China came to set the world's price floor got its own article: Why China Is Winning the AI Race.
DeepSeek-V4.1-Flash: the model that replaced the flagship-killer (and now has eyes)
On September 10, 2026 at 04:00 UTC, DeepSeek released DeepSeek-V4.1-Flash, the first model in a new architecture family — and retired V4-Flash-0731 and V4-Flash-Vision-Exp on the API. The new canonical model name is deepseek-flash; the old names (deepseek-v4-flash, deepseek-v4-flash-vision-exp) still work for compatibility but serve V4.1-Flash and bill at Flash prices. DeepSeek says "tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime" — and cut Flash prices 9–57% at the same time. Everything below the spec sheet describes the retired V4-Flash-0731 and is kept as historical context.
Every so often a "small" model comes along that makes the flagship look overpriced. On July 31, 2026, DeepSeek released the production version of V4-Flash and posted the weights on Hugging Face under the MIT license — they promptly hit the top of the trending charts. The 0731 in the name is just the release date: July 31. Independent testers found it beating DeepSeek's own flagship V4-Pro on several agentic benchmarks — at a fraction of the cost.
On September 10, 2026, DeepSeek went further and replaced the Flash tier outright with V4.1-Flash: a 552B-parameter causal encoder-decoder MoE (8B active on prefill, 16B on decode) with native image understanding — the first multimodal model DeepSeek has shipped on the API — a 1M-token context, 384K max output, and roughly a quarter of the previous KV-cache footprint. The weights are MIT-licensed on Hugging Face. The old Flash's fine-tune story (below) explains how a small model can outrun a flagship; V4.1-Flash is the same thesis with a new architecture and eyes.
The V4.1-Flash spec sheet (current)
| Specification | DeepSeek-V4.1-Flash (deepseek-flash) |
|---|---|
| Model type | Sparse Mixture-of-Experts (MoE), causal encoder-decoder (20+20 layers, 384 routed experts + 1 shared, 6 active) |
| Total parameters | 552B |
| Active parameters per token | 8B (prefill) / 16B (decode) |
| Context window | 1,000,000 tokens (1M) |
| Max output tokens | 384K |
| Input modalities | Text + native image understanding (multimodal); images are converted to tokens and billed with text input |
| License | MIT — weights free for commercial and non-commercial use |
| API model name | deepseek-flash — legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp still resolve but route here |
| Concurrency limit | 2,500 concurrent requests |
| Efficiency | ~1/4 the KV-cache footprint of the previous V4-Flash, DSpark speculative decoding built in |
A million tokens of context, MIT license, native vision, and prices that undercut most of the market. Only 8B–16B of the 552B parameters wake up for any given token — like a 552,000-person company where a few thousand people touch any one project. That's what makes it cheap to run.
The retired V4-Flash-0731 spec sheet (historical)
| Specification | DeepSeek-V4-Flash-0731 (retired Sep 10) |
|---|---|
| Model type | Sparse Mixture-of-Experts (MoE) transformer |
| Total parameters | 284B (304B with the attached DSpark speculative-decoding draft module) |
| Active parameters per token | 13B |
| Context window | 1,000,000 tokens (1M) |
| Max output tokens | 384K (~122.7 tokens/sec generation) |
| License | MIT — weights free for commercial and non-commercial use |
| Modes | Thinking (low / high / max reasoning effort) + non-thinking |
| Tool calls / context caching | Both supported |
| Concurrency limit | 2,500 concurrent requests (high-throughput friendly) |
| Self-hosting | 3-bit quantized version runs in ~110 GB of memory; full FP8 weights ≈ 166.9 GB |
How it handles a million tokens without the bill exploding
The expensive part of long-context AI is the key-value (KV) cache — attention's working memory, which grows with every token you feed in. V4 attacks it from two directions. Some attention layers keep detailed notes, condensing every 4 tokens into one entry and only re-reading the relevant ones (CSA, Compressed Sparse Attention). Other layers keep summaries, condensing every 128 tokens into one entry that always gets read (HCA, Heavily Compressed Attention).
The payoff, per Goldman Sachs research: at a full 1M-token input, V4-Flash needs only ~10% of the compute and ~7% of the KV-cache memory of DeepSeek V3.2. A few more tricks stack on top:
- Speculative decoding via the attached
DSparkdraft module — a small "intern" model drafts several tokens ahead, and the main model checks them all in one pass instead of one at a time. vLLM and SGLang enable it with a single flag:--speculative-config. - FP4 quantization for the MoE expert weights and FP8 for everything else, cutting KV-cache memory by up to 90%.
- A packaging change: Jinja chat templates are gone, replaced by a Python script package (
encoding_dsv4) and a configurablereasoning_effortparameter (low / high / max).
How it was trained (the interesting bit)
Flash-0731 was pretrained on more than 32 trillion tokens, then fine-tuned in two stages. First, DeepSeek built a separate specialist model per domain — math, coding, agentic tasks — using supervised fine-tuning plus GRPO reinforcement learning. Then it merged the 10+ specialists via on-policy distillation: the merged model wrote its own answers, and training nudged each one toward how the best specialist would have answered. That's how one model inherits ten skills without averaging them into mush.
Two details worth knowing. The three reasoning levels (low / high / max) were trained as genuinely distinct behaviors with different length penalties and context windows — "max" prepends a system prompt pushing the model to fully decompose problems and test edge cases. And during tool-using agentic tasks, Flash keeps its entire reasoning history in context across every round, including across user messages — something V3.2 simply threw away.
What independent benchmarks say (August 2026)
| Benchmark | V4-Flash-0731 (max reasoning) | Context |
|---|---|---|
| Artificial Analysis Intelligence Index | 50 (preview: 40, V4-Pro: 44) | Ties Gemini 3.6 Flash (50); 1 pt behind GPT-5.6 Luna & GLM-5.2 (51); leader is Kimi K3 (57) |
| GDPval-AA v2 (real-work head-to-head) | 1,558 Elo | 2nd-best open-weights model (behind Kimi K3's 1,685; ahead of GLM-5.2's 1,508) |
| Terminal-Bench 2.1 (CLI agent tasks) | 82.7% | +21 pts vs. April preview (61.8%) |
| τ³-Bench Banking (multi-turn tool use) | 31.1% | ~8 pts above preview |
| CodeArena WebDev (front-end dev) | 1,577 | 7th overall, 3rd among open-weights models |
| Cost per Intelligence-Index task | $0.03 | vs. $0.05 for the similar-intelligence GPT-5.6 Luna |
The story in one line, per DeepLearning.AI's The Batch: Flash-0731 sits on Artificial Analysis' Pareto frontier for intelligence vs. cost per task — no tracked model is both smarter and cheaper. It also landed during a brutal pricing week: OpenAI cut GPT-5.6 Luna by 80%, Google shipped Gemini 3.6 Flash, and Thinking Machines released Inkling Small. The market's center of gravity is now intelligence-per-dollar.
What it cost before the price change
At launch, the first-party API charged a flat $0.0028 per million cached-input tokens, $0.14 per million input tokens, and $0.28 per million output tokens (≈¥0.02 / ¥1 / ¥2). The August 16 tariff overhaul replaced that flat rate with the peak/off-peak schedule — and the September 10 release of V4.1-Flash cut Flash prices again, between 9% (output) and 57% (cache-hit input) below the August card. Current numbers are in the pricing section.
DeepSeek-V4-Pro-0813: the flagship, with caveats
DeepSeek initially announced that all deepseek-v4-pro requests would route to V4.1-Flash at Flash prices from September 14, 04:00 UTC until "V4.1-Pro" launches. That plan was reversed before it took effect: the API docs now state, "in response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." Pro keeps its own rates (below), and there is still no date for V4.1-Pro.
Pro is the 1.6-trillion-parameter flagship, and its August 13, 2026 "GA Release" was already the third V4-Pro drop in four months: an open preview on April 24, GA on July 19, then the quieter 0813 production build on August 13. Per DeepSeek's own change log, 0813 is mostly a serving-efficiency update — ~51.7B new parameters across four new DSpark speculative-decoding keys, not a new brain. On the API it simply serves as deepseek-v4-pro; no model-name migration needed.
The spec sheet
| Specification | DeepSeek-V4-Pro-0813 |
|---|---|
| Model type | Sparse Mixture-of-Experts (MoE) |
| Total parameters | 1.6 trillion |
| Active parameters per token | 49 billion |
| Context window | 1,000,000 tokens (1M) |
| Max output tokens | 384K |
| License | MIT, open weights on Hugging Face (full FP8 weights ≈ 892.7 GB — not single-workstation self-hostable) |
| Modes | Thinking (default) + non-thinking; reasoning effort low / high / max |
| API compatibility | OpenAI format (https://api.deepseek.com), Responses API, Anthropic API (/anthropic), JSON output, tool calls, Chat Prefix Completion (Beta), FIM Completion (Beta, non-thinking only) |
| Concurrency limit | 500 concurrent requests |
| Headline use case | Agentic coding; native Codex integration with one-click setup; "Expert Mode" in the DeepSeek app/web |
What's actually new under the hood
Two ideas carry the 1.6T model. mHC (Manifold-Constrained Hyper-Connections) keeps signal amplification under 2×, which keeps training a 1.6T model stable at only ~6.7% compute overhead. DSA (DeepSeek Sparse Attention) is the same hybrid compressed/hierarchical attention as Flash — CSA plus HCA — pushing context-cost growth closer to linear than quadratic. The result, per Goldman Sachs and DeepSeek's own figures: at the 1M-token setting, V4-Pro needs ~27% of the single-token inference compute and ~10% of the KV cache that V3.2 needed. The 0813 build then spends its ~51.7B new DSpark parameters on speculative decoding, i.e., faster and cheaper serving.
The benchmark reality check
DeepSeek's own numbers — run through DeepSeek Harness at max reasoning effort, versus the April preview — are big:
| Benchmark | V4-Pro-0813 (vendor) | Δ vs. preview |
|---|---|---|
| Terminal-Bench 2.1 | 87.9 | +15.8 |
| DeepSWE | 62.7 | +49.9 |
| CyberGym | 83.3 | +30.6 |
| Toolathlon-Verified | 74.1 | +18.2 |
| NL2Repo | 61.5 | +23.0 |
| SWE-bench Verified | >80% (official model card) | — |
Independent labs (AIToolsReview, Aug 18, 2026) tell a more nuanced story:
| Independent test | Result | Takeaway |
|---|---|---|
| Artificial Analysis Intelligence Index | 53 — #3 of ~106 tracked models (median 27) | Genuinely frontier-ranked on composite intelligence |
| CoderSera SWE-bench Verified (neutral harness) | 96.40% ±0.83 — #2 (behind Claude Opus 5's 97.00%) | Excellent real coding score at a fraction of the cost |
| CoderSera Terminal-Bench 2.1 (neutral harness) | 54.68% vs. vendor's 87.9 | 33-point vendor/neutral gap — and V4-Flash scored 67.04% on the same harness, beating Pro |
| Artificial Analysis AA-Omniscience (honesty) | 0.83 (near floor; Claude Opus 5: 37.07) | The model almost never says "I don't know" — weak calibration |
| NIST CAISI (April preview build) | 94% jailbreak compliance vs. 8% for US reference models; ~8-month capability gap vs. frontier | Most recent independent safety data; not yet re-run on 0813 |
The Terminal-Bench gap is the one to sit with: on a neutral harness, the smaller Flash scored 67.04% while Pro managed 54.68%. Treat vendor magnitudes as company-reported — the direction (large agentic gains) is credible and consistent across five benchmarks, but run your own evals before committing.
What it's like to actually use
Hands-on testing (via AIToolsReview) scored V4-Pro at 76.25% (61/80) on an 8-question practical test, with full marks on a hard math problem and a long-horizon agentic task. Front-end and one-shot UI generation are its clearest strength — Three.js scenes, app clones, physics demos. Known weaknesses: it overthinks and overengineers simple tasks, its honesty calibration is weak (see the AA-Omniscience row above), it's text-only, and there's no published agentic-safety evaluation.
Expect large agentic gains over the preview (credible across five vendor benchmarks), expect V4-Flash to beat Pro on many simple tasks, and add a verification step for factual work. That's the honest read.
Why is DeepSeek able to set the price floor the whole market now follows? That story — China's open-weights strategy, structural cost advantage, and domestic-silicon hedge — got its own article: Why China Is Winning the AI Race (2026). The rest of this guide stays focused on the models, the harness, and the prices.
DeepSeek Harness: DeepSeek's answer to Claude Code
On the same day as the Pro GA — August 13, 2026 — DeepSeek quietly shipped something arguably more interesting: DeepSeek Harness (dsh), an MIT-licensed, plugin-first agent framework, on GitHub at deepseek-ai/deepseek-harness. Its README states the organizing idea in three words: "everything is a plugin." DeepSeek used it internally to validate V4-Flash-0731 and to produce V4-Pro's benchmark table. Everyone calls it an open-source Claude Code rival — but the deliberate difference is that it's not locked to DeepSeek models at all.
v0.1.0-rc.5 at time of writing. The README warns plainly: there will be compatibility-breaking changes. Not yet a stable production product.
Everything is a plugin — literally
The runtime is built on Cordis, a vendored dependency-injection framework (cordiverse/cordis, designed per A Programming Paradigm for Spatiotemporal Composability). Every part of the product is a plugin — the model adapter, tool registry, session log, agent loop, sandbox, even the web UI — so every part is replaceable from configuration. There's no privileged core to patch: extensions mount beside the other plugins, and registrations are reversible effects that unwind when a plugin unloads.
The docs describe eight independently swappable layers:
ctx.llm — pluggable connection to any configured providerSessionEvent log with persistence, replay, fork, resumeIt speaks every provider
Despite the branding, the harness supports DeepSeek, Anthropic, OpenAI, Amazon Bedrock, Google Vertex, Azure, and any custom OpenAI-compatible endpoint as interchangeable inference providers. Swapping providers is a configuration change, not a rewrite — which makes dsh infrastructure a team can standardize on regardless of which model wins a given benchmark cycle.
The features that matter
- Capability seams — each swappable capability has three roles (Service Definition / Provider / Consumer). One provider swap can move Bash, PTY, LSP, and subprocess execution together to a remote sandbox, with no provider forks.
- The session log is the source of truth — "model-visible ⟺ logged": anything that reaches a model request must be reconstructable from the append-only log. Fork, resume, transcripts, telemetry, and persistence all derive from that stream.
- Typed event system — durable session events (
session/event,turn/*,step/*,tool/*) plus live extension events (agent/*,tools/*) with waterfall semantics, letting plugins intercept, rewrite, or reject requests. - Agent loop — steps (one model request + tool calls) compose into turns;
agent/pre-step,agent/request, andllm/streamare the interception points. - Sandboxing & approvals — filesystem, shell, subprocess, terminal, and sandbox are separate configurable plugins, and an approval layer gates privileged operations under the active permission policy (the Web UI asks before doing anything spicy).
- Goals, subagents, workflow, plan mode — same-session objectives (
ctx.goals), subagent delegation with pluggable providers, workflow orchestration, and plan mode as logged state. - Skills & hooks — a skill registry/provider plus Claude Code and Codex hook bridges.
- Self-modification — the agent can inspect and mount its own plugins (the
demo:cordisdemo modifies its live runtime). - Python SDK —
deepseek-harness-sdkdrives the bundled runtime as a subprocess over newline-delimited JSON-RPC on stdio.
Ways to run it
- Web UI —
npx @deepseek-ai/dsh web(orpnpm dsh webfrom source) serves a browser agent interface at http://127.0.0.1:3080. - Headless —
dsh --profile headless "task"runs one fresh persisted session and prints the final answer. - ACP — an automation-only Agent Client Protocol (JSON-RPC stdio) server.
- Plugin management —
dsh plugin --profile <name> <pnpm args>installs out-of-tree plugins per profile. - Reported operational modes: standard (general tasks), code-focused (multi-app automation), creative (custom tools), and minimal (isolated testing — reportedly how DeepSeek tested V4-Flash before release).
Installing DeepSeek Harness from source on macOS
Two ways in: npx @deepseek-ai/dsh web if you just want to look around, or from source if you want to hack on it. Here's the from-source route on a Mac.
Prerequisites
| Requirement | Version / how to get it |
|---|---|
| Node.js | ^22.19.0 or >=24.0.0 (engines field). Install via nodejs.org installer or brew install node@24 |
| pnpm | 11.7.0 (repo-pinned). With Node installed, enable Corepack: corepack enable (then pnpm --version should resolve; if not, corepack prepare pnpm@11.7.0 --activate) |
| Git | 2.26+ (brew install git if needed; Xcode Command Line Tools optional) |
| DeepSeek API key | Optional for the UI to do real work — get one at platform.deepseek.com and set DEEPSEEK_API_KEY or a repo-root .env |
| Disk space | A few GB (workspace deps + built artifacts); V4 model weights are not required (API-driven) |
The repo's native/landlock-run workspace (a Landlock self-restrict-then-exec launcher) is Linux-only sandboxing; on macOS it's simply not used, and the harness falls back to the configured non-Landlock providers. No special native toolchain is required for the standard build.
The four commands
# 1. Clone the repository
git clone https://github.com/deepseek-ai/deepseek-harness.git
cd deepseek-harness
# 2. Install dependencies (pnpm workspace; postinstall sets up lefthook git hooks)
pnpm install
# 3. Build everything (host lib -> client lib -> web frontend)
pnpm run build
# 4. Launch the Web UI (serves at http://127.0.0.1:3080 by default)
pnpm dsh webThat's it. Open http://127.0.0.1:3080, head to Settings → Models, and add your DeepSeek API key — the model route becomes usable immediately, no server restart. Then choose a workspace (your project directory), and the session composer unlocks.
Configuration & useful commands
# One-shot headless task from source (needs DEEPSEEK_API_KEY)
pnpm dsh --profile headless "summarize this workspace"
# Inspect the composed plugin tree without booting
pnpm dsh --profile web --dump-config
# Frontend dev with hot reload (HMR) — requires the dev:web watcher running
pnpm run dev:web
# Credentials live in a gitignored .env at the repo root
# DEEPSEEK_API_KEY=sk-...
# DEEPSEEK_BASE_URL=https://... # optional, defaults to public APIWhen it misbehaves
| Problem | Fix |
|---|---|
pnpm: command not found | corepack enable after installing Node; or npm install -g pnpm@11.7.0 |
corepack not present | Older Node bundles it; otherwise install Node 22.19+ fresh |
Build fails on tsdown/tsc OOM | Increase Node heap: export NODE_OPTIONS=--max-old-space-size=8192 |
| Lefthook hook errors after clone | node scripts/install-lefthook.mjs (also needed if postinstall was skipped) |
| Web UI loads but no model works | Add API key in Settings → Models; check .env for DEEPSEEK_API_KEY |
| "Port 3080 in use" | pnpm dsh web --port 8080 (app flags come after the launcher's) |
| You want the packaged binary instead | npx @deepseek-ai/dsh web — same UI, no build step |
DeepSeek API discount windows, off-peak times, and prices
On August 13, 2026, DeepSeek announced something unusual for an LLM API: peak/off-peak (峰谷) tariffing — tokens priced like electricity. It took effect at 16:00 UTC on August 16, 2026 (00:00 Beijing time, August 17), and on August 24 the fine print changed again: peak hours now apply Monday through Friday only, with weekends fully off-peak. The September 10 release of V4.1-Flash then cut Flash prices while keeping the schedule. This is the money section.
The headline answer
Every hour outside the two weekday peak windows is billed at 50% of the peak price. The peak windows are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday — and all of Saturday and Sunday is off-peak, 24 hours a day.
- Peak #1 01:00–04:00 UTC = 09:00–12:00 Beijing
- Peak #2 06:00–10:00 UTC = 14:00–18:00 Beijing
- Weekdays only peak applies Mon–Fri — weekends are fully off-peak
- Off-peak every other weekday hour — 17 hours/day, plus all weekend at 50% of peak
The purple marker shows where you are right now — it moves with live UTC time so you can see at a glance whether the API is billing you full price or half.
| Window | UTC | Beijing (UTC+8) |
|---|---|---|
| Peak #1 (full price) | 01:00–04:00 | 09:00–12:00 |
| Peak #2 (full price) | 06:00–10:00 | 14:00–18:00 |
| Off-peak (50% of peak) | every other weekday hour | every other weekday hour |
| Weekend (off-peak all day) | Sat + Sun, 24h | Sat + Sun, 24h |
- Total peak time is 7 hours per weekday, 35 hours per week — about 21% of the 168-hour week (it was 49 hours/week before the August 24 change). Everything else — 133 hours a week — is off-peak. Official DeepSeek docs: "Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday (all other hours are off-peak)."
- Practical read for the US: on Eastern time (summer), peak runs 21:00–00:00 and 02:00–06:00 — your 9-to-5 workday bills off-peak, and the overnight batch job someone scheduled "because nobody is using anything then" is the one that pays double. When US clocks fall back on November 1 (October 25 in the EU), each local window shifts one hour earlier. For China-based teams, lunch (12:00–14:00) and everything after 18:00 Beijing is off-peak.
Current per-1M-token prices (official, checked Sep 28, 2026)
deepseek-flash (V4.1-Flash) — concurrency 2,500. Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp route here and bill at these rates. Image inputs are converted to tokens and billed with text input:
| Token type | Off-peak 50% | Peak 100% |
|---|---|---|
| Input — cache hit | $0.003 ≈¥0.02 | $0.006 ≈¥0.04 |
| Input — cache miss | $0.15 ≈¥1.00 | $0.30 ≈¥2.00 |
| Output | $0.60 ≈¥4.00 | $1.20 ≈¥8.00 |
deepseek-v4-pro (Pro-0813) — concurrency 500, billing unchanged after the September 14 reroute was reversed:
| Token type | Off-peak 50% | Peak 100% |
|---|---|---|
| Input — cache hit | $0.022 ≈¥0.15 | $0.044 ≈¥0.30 |
| Input — cache miss | $0.66 ≈¥4.50 | $1.32 ≈¥9.00 |
| Output | $1.98 ≈¥13.50 | $3.96 ≈¥27.00 |
Off-peak is exactly half of peak in every category. Base URLs: https://api.deepseek.com for OpenAI format, https://api.deepseek.com/anthropic for Anthropic format. Billing deducts from granted balance first, then topped-up balance.
What changed on September 10 (the price cut)
With V4.1-Flash, DeepSeek cut Flash rates between 9% and 57% versus the August card. The cache-hit cut is the steepest — the same stable-prefix workloads got dramatically cheaper:
| Token category | V4-Flash (Aug 21 rates) | V4.1-Flash (Sep 10 rates) | Cut |
|---|---|---|---|
| Cache-hit input (off-peak) | $0.007 | $0.003 | −57% |
| Cache-miss input (off-peak) | $0.22 | $0.15 | −32% |
| Output (off-peak) | $0.66 | $0.60 | −9% |
The honest part: base prices are still above the original flat rate
The original July flat rates for V4-Flash were ¥0.02 / ¥1 / ¥2 (cache-hit-in / cache-miss-in / output) — about $0.0028 / $0.14 / $0.28. The new V4.1-Flash card is much closer to that floor than the August card was, but it hasn't fully returned to it:
deepseek-flash (V4.1-Flash) — increase over the original flat rate:
| Token category | Peak | Off-peak |
|---|---|---|
| Cache-hit input | +114% | +7% |
| Cache-miss input | +114% | +7% |
| Output | +329% | +114% |
deepseek-v4-pro — increase over its GA flat rate (unchanged since August 16):
| Token category | Peak | Off-peak |
|---|---|---|
| Cache-hit input | +1,100% | +500% |
| Cache-miss input | +200% | +50% |
| Output | +350% | +125% |
"Off-peak" is 50% off the peak price — not a discount versus the original flat rate. For V4.1-Flash, off-peak input is back within ~7% of the original flat rate, but off-peak output still costs ~2.14× it, and peak output ~4.29×. Pro's off-peak output runs ~2.28× its original flat rate, and peak ~4.55×. Plan your cost model on the current table, not the "discount" label.
Three real levers to pull
- Shift work to off-peak (50% off — and weekends are all off-peak). Best for batch data processing, bulk translation, data cleaning, offline analysis, overnight model evaluation, report generation — anything with loose latency requirements. Interactive chat, online agents, and real-time copilots generally can't wait. With the August 24 change, a weekend batch pays the off-peak rate all 48 hours.
- Maximize cache hits (the biggest multiplier). Cache-hit input is ~50× cheaper than cache-miss for Flash ($0.003 vs. $0.15), ~30× for Pro ($0.022 vs. $0.66). Reuse stable system prompts, long shared contexts, and prefix-stable requests so more of your input tokens land on the cache — and note the September 10 cut hit cache prices hardest (−57%).
- Pick the right tier for the job. Flash (deepseek-flash) has 5× the concurrency (2,500 vs. 500), ~3.3× cheaper output, native vision, and — per DeepSeek's own testing — beats Pro on performance, cost, speed and total runtime. Route high-throughput, multimodal, or non-thinking workloads to Flash, and reserve Pro for the complex agentic and front-end work where the 1.6T model's depth still earns its premium.
Why DeepSeek did this
Model inference occupies AI compute around the clock, but traffic concentrates in working hours while idle periods under-utilize the clusters. Pricing tokens like electricity (峰谷电价) uses the price lever to smooth demand, raise utilization, and relieve peak pressure — a sign the LLM price war is entering a "2.0" phase of fine-grained compute-scheduling economics (智东西/Zhidx, Techritual).
Frequently asked questions
Since September 10, 2026, the Flash tier is DeepSeek-V4.1-Flash: a 552B-parameter MoE with native image understanding, ~$0.15/$0.60 per million tokens off-peak, 2,500 concurrent requests, and — per DeepSeek's own testing — it beats V4-Pro on performance, cost, speed and total runtime. Pro (1.6T total, 49B active) is the flagship for complex agentic coding and front-end generation, with 500 concurrent requests, text-only input, and unchanged billing after the September 14 reroute was reversed. V4.1-Pro is expected eventually; there is no date.
Off-peak pricing — 50% of the peak rate — applies Monday through Friday except the two peak windows (01:00–04:00 and 06:00–10:00 UTC, i.e. 09:00–12:00 and 14:00–18:00 Beijing time), and all of Saturday and Sunday is off-peak. That is 7 peak hours per weekday — 35 of 168 hours per week — and 133 off-peak hours.
It depends which "before." Versus the September 10 price cut: off-peak is exactly 50% of the current peak price, and the September cut lowered Flash rates 9–57% versus the August card. Versus the original July flat rate: V4.1-Flash off-peak input is back within ~7% of it, but output still costs ~2.14× off-peak and ~4.29× at peak. Plan on the current table, not the "discount" label.
The two legacy model names were fully retired on July 24, 2026 at 15:59 UTC — requests using them now return errors. Migrate to model: "deepseek-flash" (thinking mode is a request parameter, not a separate model name). The V4-era names deepseek-v4-flash and deepseek-v4-flash-vision-exp also route to V4.1-Flash and bill at Flash prices, but pin deepseek-flash going forward.
The harness is MIT-licensed and free — you only pay for whichever inference provider you plug in. It supports DeepSeek, Anthropic, OpenAI, Amazon Bedrock, Google Vertex, Azure, and any custom OpenAI-compatible endpoint.
git clone https://github.com/deepseek-ai/deepseek-harness.git && cd deepseek-harness && pnpm install && pnpm run build && pnpm dsh web. You will need Node 22.19+/24+, pnpm 11.7.0 (via Corepack), and Git 2.26+. The web UI serves at http://127.0.0.1:3080 by default.
DeepSeek says V4.1-Flash beats V4-Pro on performance, cost and speed, and independent tests of the retired V4-Flash already showed Flash beating Pro on a neutral Terminal-Bench harness (67.04% vs. 54.68%). Practical rule: start with deepseek-flash for throughput and cost; keep Pro for the hardest multi-step tasks until V4.1-Pro ships.
Since September 10, 2026, deepseek-flash (V4.1-Flash) supports native image understanding — images are converted to tokens and billed as input. V4-Pro remains text-only.
Yes. V4.1-Flash and V4-Pro-0813 are both MIT-licensed open weights on Hugging Face. V4.1-Flash is the practical self-host tier; Pro's full FP8 weights are ~892.7 GB — multi-GPU server territory.
Sources & further reading
Official DeepSeek
- DeepSeek-V4.1-Flash Release — DeepSeek API Docs (Sep 10, 2026)
- Models & Pricing — DeepSeek API Docs (verified Sep 28, 2026)
- Change Log — DeepSeek API Docs
- Introducing DeepSeek-V4.1-Flash — DeepSeek
- DeepSeek-V4-Pro GA Release — DeepSeek API Docs
- DeepSeek Harness on GitHub (deepseek-ai/deepseek-harness)
- DeepSeek-V4-Flash-0731 on DeepInfra (historical)
Models & benchmarks
- DeepSeek Pushes the Frontier Again (V4-Flash fine-tune) — DeepLearning.AI The Batch
- DeepSeek Open Sources Production DeepSeek-V4-Flash Under MIT Licence — Open Source For You
- DeepSeek V4 Pro 0813: Benchmarks & Verdict — AIToolsReview
- DeepSeek V4 Pro Benchmarks: Official vs Independent — OrcaRouter
- GLM-5.3 vs DeepSeek V4-Pro — Flowtivity
- 高盛:DeepSeek V4 对中国 AI 意味着什么?(Goldman Sachs compute-efficiency research) — 全天候科技
Harness
Related reading
Pricing & peak/off-peak tariff
- DeepSeek Pricing (Sep 2026): V4.1 Flash & V4 Pro API Rates — Justin McKelvey
- DeepSeek V4.1 Flash Pricing: Pro Routed From Sept 14 — TokenCost
- DeepSeek Keeps V4-Pro Alive After User Backlash — BeingGuru
- DeepSeek-V4.1-Flash Pricing Explained: Peak vs Off-Peak — Apifox/Apidog
- 最高涨 1100%!DeepSeek 新定价今日生效 — 智东西
- DeepSeek 调价今日生效:首推"峰谷定价" — 湖北日报
- DeepSeek API 推出峰谷定價方案,空閒時段費用減半 — Techritual
Last updated: September 28, 2026. Prices and availability verified against DeepSeek's official API documentation and the sources above as of September 28, 2026. Key changes since the previous update: DeepSeek-V4.1-Flash shipped September 10 (new canonical name deepseek-flash, Flash prices cut 9–57%, native vision); peak hours became weekdays-only on August 24 (weekends fully off-peak); the planned September 14 V4-Pro reroute was reversed, so Pro continues at unchanged billing; and the legacy names deepseek-chat/deepseek-reasoner were fully retired July 24. DeepSeek reserves the right to adjust pricing — re-check the official Models & Pricing page before making budget decisions. Vendor-reported benchmark figures should be treated as company-reported until independently reproduced.