Model comparison
DeepSeek-V4.1-Flash or Gemini 3.8 Flash? Price and caching
DeepSeek-V4.1-Flash costs $0.22 for a cached coding session that costs $0.71 on Gemini 3.8 Flash, and Gemini's introductory rates end December 31, 2026.
· Prices as of September 28, 2026
DeepSeek-V4.1-Flash
DeepSeek · Released September 10, 2026
DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.
DeepSeek-V4.1-Flash facts and comparisonsGemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisons
The short answer
DeepSeek-V4.1-Flash costs less on every line, and the example agentic coding session comes to $0.22 against $0.71 on Gemini 3.8 Flash, a 3.2x gap, wider than on uncached work because DeepSeek's cache hits cost $0.006 per million. Gemini 3.8 Flash is the pick if you work in Gemini CLI or want Google's free tier, and its prices here are introductory rates that rise on January 1, 2027. DeepSeek-V4.1-Flash suits cost-driven work through OpenRouter or OpenCode, or on your own hardware with its MIT-licensed weights.
Choose DeepSeek-V4.1-Flash if
- You want the lower rate on every line: $0.30 input and $1.20 output at DeepSeek's peak, against $0.75 and $3.75.
- Your agent rereads a big cached context, where DeepSeek's hit costs $0.006 per million against Gemini's $0.075, a 12.5x difference.
- You need long single responses: DeepSeek lists 384K output tokens and Gemini 3.8 Flash 65.5K.
- You want open weights under the MIT license.
Choose Gemini 3.8 Flash if
- You work in Gemini CLI, where Gemini 3.8 Flash is the Flash half of the default auto model for Gemini API key and Vertex AI users.
- You want to start on the Gemini API free tier, which covers its input, output, and caching.
- You also use Cursor or GitHub Copilot, which offer Gemini 3.8 Flash but not DeepSeek-V4.1-Flash.
- You want the model Google calls its most intelligent Flash, aimed at long-horizon software engineering and complex multi-file refactoring.
Side by side
Specs and prices
| Fact | DeepSeek-V4.1-Flash | Gemini 3.8 Flash |
|---|---|---|
| Maker | DeepSeek | |
| API model id | deepseek-flash | gemini-3.8-flash |
| Released | September 10, 2026 | September 2, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 384K tokens | 65.5K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $0.30 | $0.75 |
| Cache hit, per 1M | $0.006 | $0.075 |
| Cache write, per 1M | $0.30 (same as input) | $0.75 (same as input) |
| Output, per 1M tokens | $1.20 | $3.75 |
| Runs in | OpenCode and OpenRouter | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | DeepSeek-V4.1-Flash | Gemini 3.8 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.22 | $0.71 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.06 | $0.15 |
| Output-heavy generation, 30K input, 80K output | $0.11 | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $24.42 | $78.38 |
| Where the session’s cost goes | ||
| Cache writes | $0.12 | $0.30 |
| Cache reads | $0.01 | $0.15 |
| Uncached input | $0.03 | $0.08 |
| Output | $0.06 | $0.19 |
- caching saves on the session with DeepSeek-V4.1-Flash (73%)
- $0.59
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
Neither model charges a cache-write premium
Comparisons with Claude or GPT models often turn on cache-write fees. This one does not. DeepSeek lists no fee for writing its cache, and Google publishes no separate cache-write price, so on both models the tokens a coding agent writes to the cache are billed as ordinary input: $0.30 per million on DeepSeek-V4.1-Flash and $0.75 on Gemini 3.8 Flash.
That leaves list rates and cache hits. On uncached work the gap follows the rate cards: 2.5x on the large one-off review, $0.06 against $0.15, and 2.9x on the output-heavy generation, $0.11 against $0.32, where Gemini's $3.75 output rate weighs most.
Cache hits are where the two split. Google bills a hit at 10% of input, $0.075 per million. DeepSeek bills $0.006, 2% of its input price, and ties that to a KV cache that needs a quarter of the memory of the previous generation. In the example session the 2M cached tokens cost $0.15 on Gemini and $0.01 on DeepSeek, and the whole session costs $0.71 against $0.22, a 3.2x gap.
Gemini 3.8 Flash's introductory rates end on December 31, 2026
Google lists the $0.75 input, $0.075 cached, and $3.75 output rates as introductory, through December 31, 2026. From January 1, 2027, Gemini 3.8 Flash costs $1.50 input, $0.15 cached, and $7.50 output per million tokens. Every figure on this page uses today's rates, so if DeepSeek's prices hold, the gap in 2027 will be wider than the tables show.
DeepSeek's price list has a discount built in. The figures here are its peak rates, which apply from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. Outside those hours DeepSeek charges 50% less.
Google's caching has one more cost to know about. Implicit caching is on by default, but a hit is not certain. Explicit caching gives a named cache with an assured discount and adds a storage charge while it lives, $0.50 to $1 per million tokens per hour on Flash models. The tables include no storage charge.
Gemini CLI, OpenCode, and the free tier
Gemini CLI is Google's own coding agent and runs Gemini models. For Gemini API key and Vertex AI users, its default auto model uses Gemini 3.8 Flash as the Flash half, and Gemini CLI sends a high thinking level where the API default is medium. Gemini 3.8 Flash is also in Cursor, OpenRouter, OpenCode, and GitHub Copilot.
DeepSeek-V4.1-Flash is not in Cursor or Copilot. It runs through OpenRouter and OpenCode, and on DeepSeek's API as deepseek-flash. Its MIT-licensed weights allow self-hosting too, at a hardware cost this page does not estimate. Through OpenRouter, requests for it go to one of several providers, and their prices can differ from DeepSeek's own.
Google calls Gemini 3.8 Flash "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows," and claims better robustness against prompt injection. DeepSeek calls V4.1 Flash the smallest model in its new architecture family, with native visual understanding. Output limits differ sharply: 65.5K per response on Gemini, 384K on DeepSeek.
To compare them on your own work, EveryToken prices Gemini 3.8 Flash at Google's rates from Gemini CLI, Cursor, or OpenCode history, and prices DeepSeek-V4.1-Flash when you use it through OpenRouter.
Prompt caching
How each maker bills cached tokens
DeepSeek
DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.
A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.
Source: DeepSeek API: Context caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what DeepSeek-V4.1-Flash and Gemini 3.8 Flash really cost you.
everyaitoken reads your OpenRouter, Gemini CLI, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is DeepSeek-V4.1-Flash cheaper than Gemini 3.8 Flash?
Yes, on every line at current rates. The example cached session costs $0.22 against $0.71, and 110 sessions a month come to $24.42 against $78.38. The gap grows on cache-heavy work because DeepSeek's cache hits cost $0.006 per million against $0.075.
When do Gemini 3.8 Flash prices go up?
Google's introductory rates run through December 31, 2026. From January 1, 2027, the model costs $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Does either model charge for cache writes?
No. DeepSeek lists no fee for writing its cache, and Google publishes no separate write price, so written tokens cost ordinary input on both. Google's explicit caching adds a storage charge for as long as the cache lives.
Which model allows longer outputs?
DeepSeek-V4.1-Flash lists up to 384K output tokens per response, and Gemini 3.8 Flash 65.5K. Both accept about 1M tokens of context: 1M on DeepSeek and 1.05M on Gemini.
Sources
- DeepSeek API: Models and pricing
- DeepSeek: V4.1 Flash release
- DeepSeek API: Change log
- OpenRouter: DeepSeek-V4.1-Flash
- OpenRouter: Rankings
- OpenCode docs: Zen
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- DeepSeek API: Context caching
- Google: Context caching