DeepSeek API Price Increase: Up to 12x, Peak Hours (2026)
DeepSeek's new API prices went live 2026-08-16 16:00 UTC. Cache hits jumped 12x on V4 Pro, output 4.7x on Flash. Peak hours, new rates, and the bill math.
DeepSeek’s new API prices took effect at 16:00 UTC on 2026-08-16, and they are not a flat increase. The V4 family moved to peak and off-peak billing, so the same call now costs one of two prices depending on the hour.
Every tier of the new schedule sits above the old flat rate. The tier that went up most is the one that used to be nearly free.
Effective: 2026-08-16, 16:00 UTC (= 2026-08-17, 00:00 Beijing time)
Billing model: peak / off-peak, off-peak is half of peak
Peak hours: 01:00-04:00 and 06:00-10:00 UTC, 7 hours a day
V4 Flash peak: $0.014 hit / $0.44 miss / $1.32 out, per 1M
V4 Flash off-peak: $0.007 / $0.22 / $0.66
V4 Pro peak: $0.044 / $1.32 / $3.96
V4 Pro off-peak: $0.022 / $0.66 / $1.98
Biggest jump: V4 Pro cache hit, 12.1x at peak
Smallest jump: V4 Pro cache miss, 1.5x off-peak
Unchanged: 1M context, 384K max output, concurrency 2,500 / 500
What Changed in DeepSeek’s API Pricing?
DeepSeek replaced one flat price per tier with two, and raised both. The mechanism is new, not just the numbers. Before 2026-08-16 a cache miss on deepseek/deepseek-v4-flash cost $0.14 per million tokens at any hour of any day. It now costs $0.22 or $0.44 depending on when the request lands.
Here is the full table as published on the official pricing page, against the rates that were live until the change.
| Tier, per 1M tokens | Old flat | New off-peak | New peak |
|---|---|---|---|
| V4 Flash cache hit | $0.0028 | $0.007 | $0.014 |
| V4 Flash cache miss | $0.14 | $0.22 | $0.44 |
| V4 Flash output | $0.28 | $0.66 | $1.32 |
| V4 Pro cache hit | $0.003625 | $0.022 | $0.044 |
| V4 Pro cache miss | $0.435 | $0.66 | $1.32 |
| V4 Pro output | $0.87 | $1.98 | $3.96 |
Old rates are our own snapshots of the same page, taken 2026-08-06 for Flash and 2026-08-13 for Pro, and published at the time in 6 ways to pay less as prices rise and V4 Pro 0813.
Nothing else on the page moved. Model IDs are still deepseek-v4-flash and deepseek-v4-pro with no suffixes, all three protocols work on both models, and the deduction rule is unchanged: tokens times price, granted balance first.
One thing did disappear. Since 2026-08-06 the page carried a footnote warning that a significant increase was coming, with no amount and no date. That footnote is gone, replaced by the one explaining off-peak hours.
When Did the New DeepSeek Prices Take Effect?
16:00 UTC on 2026-08-16, which is 00:00 Beijing time on 2026-08-17. Both are the same instant. If you have seen two different dates quoted, that is why.
The English change log says the new prices “take effect at 16:00 (UTC Time) on August 16, 2026.” The Chinese change log for the same entry says 北京时间 2026 年 8 月 17 日 0 时, Beijing time, midnight on the 17th. Beijing is UTC+8, so 00:00 on the 17th there is 16:00 on the 16th in UTC.
This tripped people up in practice. A post on r/DeepSeek on 2026-08-15 asked why the model already felt more expensive “if new prices go in live from August 17,” quoting the Chinese notice from a timezone where it was still the 15th.
The announcement landed three days ahead of the change, in the 2026-08-13 change log entry that also shipped V4 Pro GA.
What Are DeepSeek’s Peak Hours in Your Time Zone?
Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, seven hours a day, and whether that hurts depends entirely on where you sit. DeepSeek publishes the windows in UTC on the English page and in Beijing time on the Chinese one. Neither is the timezone most readers work in.
| Local zone (Aug 2026) | Peak window 1 | Peak window 2 | Peak hours inside a 09:00-18:00 workday |
|---|---|---|---|
| Beijing, Singapore (UTC+8) | 09:00-12:00 | 14:00-18:00 | 7 of 9 |
| Tokyo, Seoul (UTC+9) | 10:00-13:00 | 15:00-19:00 | 6 of 9 |
| India (UTC+5:30) | 06:30-09:30 | 11:30-15:30 | 4.5 of 9 |
| Berlin, Paris (UTC+2, CEST) | 03:00-06:00 | 08:00-12:00 | 3 of 9 |
| London (UTC+1, BST) | 02:00-05:00 | 07:00-11:00 | 2 of 9 |
| New York (UTC-4, EDT) | 21:00-00:00 prev day | 02:00-06:00 | 0 of 9 |
| San Francisco (UTC-7, PDT) | 18:00-21:00 prev day | 23:00-03:00 | 0 of 9 |
The windows track a Chinese working day, which is what you would expect from a load-shaping measure. A team in Shenzhen pays peak for seven of nine working hours. A team in California pays peak for none of them, unless it runs overnight batches.
One r/DeepSeek user in Pacific time worked the same conversion out on 2026-08-14, under the title “Sometimes being in PST rocks.”
Two caveats before you schedule around this:
- The offsets above are northern-hemisphere summer time. London and Berlin each gain an hour of peak exposure when the clocks go back. The US windows stay clear of the workday either way.
- DeepSeek has not documented which timestamp it bills by. A session that starts at 09:55 UTC and runs past 10:00 is an open question. Measure it rather than assume it.
How Much Did DeepSeek Actually Raise Prices?
Between 1.5x and 12.1x, depending on the model, the tier and the hour. There is no single multiplier, which is why the numbers people quote range from “50% more” to “10x” and both camps are right about their own workload.
| Tier | Off-peak vs old | Peak vs old |
|---|---|---|
| V4 Flash cache hit | 2.5x | 5.0x |
| V4 Flash cache miss | 1.6x | 3.1x |
| V4 Flash output | 2.4x | 4.7x |
| V4 Pro cache hit | 6.1x | 12.1x |
| V4 Pro cache miss | 1.5x | 3.0x |
| V4 Pro output | 2.3x | 4.6x |
The 12.1x is V4 Pro cache hits, from $0.003625 to $0.044 per million at peak. That is the single largest move in the table, and it is why the loudest argument on r/DeepSeek that week, posted 2026-08-14 and still sitting at a contested 0.69 upvote ratio, was about cache pricing rather than output pricing.
Cache misses, the tier that already cost the most, went up the least. Cache hits, the tier that was close to free, went up the most. Your own multiplier is a function of your cache hit rate, and it moves in the direction most people do not expect.
Why Did Cache Hits Go Up the Most?
Because the old cache hit price was the outlier, not the new one. At $0.0028 per million, a V4 Flash cache hit was 50x cheaper than a miss. At $0.014 it is 31x cheaper. Both ratios are aggressive; the old one was extraordinary.
That ratio was the whole basis of DeepSeek’s cost story for agent work. Agent loops resend the same system prompt and the same file contents every turn, so hit rates above 90% are normal. Under the old table, a 90%-cached workload paid almost nothing for its input.
So the higher your cache hit rate today, the larger your increase. Same model, same volumes, two hit rates:
- 40% hit rate, mostly fresh context: 1.7x off-peak, 3.4x at peak
- 90% hit rate, typical agent loop: 2.1x off-peak, 4.1x at peak
Caching is still worth having, and the prompt caching setup guide applies unchanged, since DeepSeek’s cache is automatic. It is just no longer the lever that makes the bill disappear.
Is Off-Peak a Discount or a Smaller Increase?
A smaller increase. The pricing page footnote reads “Off-peak rates are half of the peak rates,” which is accurate and reads like a discount. It is a comparison against the new peak price, not against what you paid last week.
Against the old flat rate, off-peak V4 Flash output is $0.66 versus $0.28. That is 2.4x more expensive during the cheap hours. There is no hour of the day at which any tier of either model costs what it cost on 2026-08-15.
The wording matters because DeepSeek used this same mechanism as a real discount before, covered further down. Off-peak in 2025 meant paying under list. Off-peak in 2026 means paying half of a list price that went up first.
What Does the Increase Do to a Real Monthly Bill?
Roughly 2x off-peak and 4x at peak for a typical agent workload. Below is one month on V4 Flash: 300M input tokens at a 90% cache hit rate, 20M output tokens. That shape is ordinary for a coding agent running most working days.
| Scenario | Cache hits | Cache misses | Output | Total |
|---|---|---|---|---|
| Old flat rate | $0.76 | $4.20 | $5.60 | $10.56 |
| New, all off-peak | $1.89 | $6.60 | $13.20 | $21.69 |
| New, all peak | $3.78 | $13.20 | $26.40 | $43.38 |
| DeepInfra, no time-of-day | $4.32 | $2.40 | $3.60 | $10.32 |
Three things fall out of that table:
- Output is now the dominant line, at 61% of the off-peak bill against 53% before. The tier you cannot cache is the tier you now pay for.
- The all-off-peak column is the floor for a US or European team, and it is still roughly double the old bill.
- A third-party host at flat pricing now lands under what DeepSeek itself charged before the change.
- Same month at a 40% hit rate instead of 90%: $31.14 old, $53.64 off-peak, $107.28 peak. Higher bill, lower multiplier, same table.
Are Third-Party DeepSeek Hosts Cheaper Now?
On V4 Flash, nearly all of them. On V4 Pro, none of them beat DeepSeek’s off-peak rate. The two models ended up in opposite positions, and the reason is how long each one’s weights have been public.
Checked on the OpenRouter endpoints API at 07:15 UTC on 2026-08-17, inside a peak window, V4 Flash 0731 had 28 listings:
| Host | Input | Output | Cache read | Context | Quant |
|---|---|---|---|---|---|
| DeepInfra | $0.08 | $0.18 | $0.016 | 1M | fp8 |
| DigitalOcean | $0.08 | $0.252 | $0.0252 | 1M | unknown |
| GMICloud | $0.084 | $0.168 | $0.0168 | 1M | fp8 |
| BaseTen | $0.13 | $0.26 | $0.028 | 1M | fp8 |
| Fireworks, Novita, Together, others | $0.14 | $0.28 | $0.028 | 1M | mixed |
| Cloudflare | $0.44 | $1.32 | $0.014 | 1.28M | unknown |
| DeepSeek | $0.44 | $1.32 | $0.014 | 1M | fp8 |
DeepSeek is now the joint most expensive input price on the list it used to anchor. Cloudflare matches it to the cent, which is what a resale of the source endpoint looks like. Everyone else running the open weights on their own hardware sits 3x to 5x below.
Two to three of those listings carry a degraded or offline status at any given moment, and which ones changes through the day. OpenRouter exposes that per endpoint, so check it at the moment you route rather than trusting a table anyone published earlier.
DeepSeek still wins one column on Flash: cache read, at $0.014 peak and $0.007 off-peak against DeepInfra’s $0.016. In early August the same comparison was $0.0028 against $0.018, a 6.4x edge. At 1.14x it no longer carries the rest of the table.
- Break-even against DeepInfra on input alone: 94% cache hit rate off-peak, 99.4% at peak. Below those rates the third party is cheaper. Before the change the same calculation gave 77%.
- Add any output volume at all and DeepInfra wins outright, because its output price is 3.7x under DeepSeek’s off-peak rate.
V4 Pro 0813 is the mirror image. It had nine listings the same morning, up from three on 2026-08-14, because its weights only went public on 2026-08-13. The cheapest of them, Novita at $1.056 in / $3.168 out, still sits above deepseek/deepseek-v4-pro’s off-peak $0.66 / $1.98.
Against DeepSeek’s peak rate the Pro picture flips again: Novita wins below a 53% cache hit rate, GMICloud below 64%.
Two caveats on the third-party route:
- You are not always buying the same model. Several cheap Flash endpoints run fp4 quantization or truncate context to 262K, and none replicate DeepSeek’s own caching stack.
- The slug trap from the last release still applies.
deepseek/deepseek-v4-flashwithout the-0731suffix is the April build.
Has DeepSeek Changed Its Pricing Model Before?
Three times in eighteen months, and this is the first one that used time-of-day to charge more rather than less. The mechanism is not new to DeepSeek. Its direction is.
| Date | Change | Direction |
|---|---|---|
| 2025-02-26 | Off-peak discounts introduced, 16:30-00:30 UTC daily, V3 at 50% off and R1 at 75% off | Down |
| 2025-09-05 | New pricing takes effect, off-peak discounts end | Up |
| 2026-08-16 | Peak / off-peak billing, off-peak set at half of peak | Up |
The 2025 announcement is still online, and the wording reads differently now: “Starting today, enjoy off-peak discounts on the DeepSeek API Platform from 16:30–00:30 UTC daily.”
That window falls entirely inside what the current schedule calls off-peak. The hours DeepSeek once discounted are the hours it now merely does not surcharge. Both the 2025 change and this one took effect at 16:00 UTC, which is worth knowing if you are watching for the next one.
How Do You Keep One Key When a Provider Reprices?
By pointing your tooling at an endpoint that is not a single vendor. Every path above commits a base URL, an API key and a billing relationship to one company’s pricing decisions. This week showed how fast those move.
Each route breaks differently. Going direct to DeepSeek ties you to peak windows on a Chinese working day and one top-up balance. A third-party Flash host is cheaper per token but ships a different quantization, sometimes a shorter context, and needs a separate key for every failover target. Staying put means paying whatever the next change log entry says.
All three speak OpenAI-compatible HTTP, so the fix is the same on any of them. Keep the base URL and the key constant, and make the model string the only thing that changes. On ofox that is deepseek/deepseek-v4-flash and deepseek/deepseek-v4-pro, so moving a workload off DeepSeek is a string edit rather than a migration.
Be clear about what that does and does not buy you on price. ofox carries both models at DeepSeek’s new peak rate, flat: $0.44 in / $1.32 out on Flash, $1.32 / $3.96 on Pro, cache read $0.014 and $0.044. There is no off-peak tier.
So if your team dodges the peak windows and you only ever call one model, DeepSeek direct is cheaper for those hours and you should use it. The gateway earns its place when you want to A/B Flash against a third-party host, or against something outside the DeepSeek family, without opening three accounts to find out.
If you run DeepSeek’s own agent tooling, the same substitution works there: the DeepSeek Harness setup guide covers pointing dsh at an arbitrary OpenAI-compatible provider, which is how several people in the r/DeepSeek threads moved off the official endpoint without changing tools. Worth knowing before you lean on it: dsh ships no per-turn cost display, so the hour your requests land in stays invisible from inside the tool. Its version and update status covers what else to check first.
Three pages go deeper than this one on the decisions that follow from the new table: six ways to pay less on the levers, V4 Pro vs Flash on which model, and the DeepSeek API pricing guide on account setup and keys.
References
- DeepSeek models and pricing
- DeepSeek pricing, Chinese edition
- DeepSeek change log
- DeepSeek context caching documentation
- DeepSeek off-peak discount announcement, 2025-02-26
- DeepSeek pricing change announcement, 2025-08-21
- DeepSeek V4 Pro GA announcement, 2026-08-13
- OpenRouter: deepseek/deepseek-v4-flash-0731 endpoints
- OpenRouter: deepseek/deepseek-v4-pro-0813 endpoints
- r/DeepSeek: New prices activated, 2026-08-16
Frequently Asked Questions
- Is the DeepSeek price increase the same for V4 Flash and V4 Pro?
- No. The two models moved by different multiples on every tier. V4 Flash cache hits went from $0.0028 to $0.014 at peak, 5x. V4 Pro cache hits went from $0.003625 to $0.044 at peak, 12.1x. Cache misses moved least on both, about 3x at peak. Output landed between 4.5x and 4.7x on both.
- Does DeepSeek bill peak rates by when a request starts or when it finishes?
- DeepSeek has not published this. The pricing page states the peak windows and the deduction rule (tokens times price) and says nothing about which timestamp applies to a request that straddles a boundary. If you run long agent sessions across 04:00 or 10:00 UTC, treat the boundary as unknown and measure your own invoices rather than assuming.
- Does the increase apply to the Anthropic-format endpoint too?
- Yes. The prices are per model, not per protocol. Both api.deepseek.com and api.deepseek.com/anthropic bill the same table, and the Responses API format is now supported on both models as well.
- Are third-party DeepSeek hosts cheaper than DeepSeek now?
- On V4 Flash, almost all of them. Checked on the OpenRouter endpoints API on 2026-08-17, 26 of 28 listings quoted a lower input price than DeepSeek's own peak rate of $0.44, with DeepInfra at $0.08 in / $0.18 out. The one match at $0.44 is Cloudflare, which mirrors DeepSeek to the cent. On V4 Pro it is the reverse: DeepSeek's off-peak $0.66 undercuts every one of the nine listings.
- Is prompt caching still worth using after the increase?
- Yes, but it saves less than it did. A V4 Flash cache hit is still about 31x cheaper than a miss at the same hour ($0.014 against $0.44 at peak). The old ratio was 50x. Caching went from the dominant lever on your bill to one of several.
- Did DeepSeek grandfather old prices for balances topped up before the change?
- There is no such statement in the change log or on the pricing page. The deduction rule is unchanged: expense equals tokens times the price in effect, drawn from granted balance first. Assume your existing balance buys tokens at the new rates.
- What happened to the old cache hit price of $0.0028?
- It is gone. V4 Flash cache hits are $0.007 off-peak and $0.014 at peak, so even the cheaper half of the new schedule is 2.5x the old flat rate. Any page still quoting $0.0028 per million for cache hits is a snapshot taken before 2026-08-16, whoever published it.


