# everyaitoken > everyaitoken is a native macOS app for Codex, Claude Code, Cursor, OpenRouter, Gemini CLI, and OpenCode. See limits with exact reset times, usage and API-equivalent cost by tool and model, and what prompt caching saved or cost, in your menu bar and 17 desktop widgets. $9 once. Canonical website: https://everyaitoken.com/ Product reference: https://everyaitoken.com/product ## Product facts - **Product:** everyaitoken is a native macOS app for tracking AI coding usage, supported account limits, and API-equivalent costs. - **Supported tools:** Codex, Claude Code, Cursor, OpenRouter, Gemini CLI, OpenCode. - **Price:** $9 once for one Mac. No subscription. Checkout is not yet available on this website. - **System requirements:** macOS 14+ · Apple silicon. - **Limit coverage:** Session and weekly limit observations are available for Claude Code and Codex. Other tools have different usage and spend coverage. - **Cost accounting:** API-equivalent estimates price input, cache reads, cache writes, and output separately. Provider-reported spend is shown separately. Subscription charges are not inferred from token estimates. - **Cache savings:** Savings compare equivalent ordinary input with cache reads and writes, including write premiums. Savings can be negative. - **History:** Reads retained local history and supported provider data. Deleted source history cannot be recovered. Missing prices or incomplete sources are labeled partial. - **Privacy:** Usage history stays locally on the Mac. Prompts, responses, and code are not retained. Credentials use macOS Keychain. Optional PostHog product analytics require opt-in and exclude usage history and credentials. - **Widgets and export:** 17 native macOS widgets, a menu bar view, multiple accounts, date filters, and CSV export. Widgets use the app's latest summary; macOS controls widget refresh timing. - **Scope:** Tracks Codex rather than ChatGPT conversations or the OpenAI API platform dashboard. It is not an invoice or a billing authority. ## Supported tools and connection sources - Codex: Local history. Limits from Codex itself. - Claude Code: Local history. Limits from Claude Code’s status line. - Cursor: Its saved sign-in, or an Enterprise admin API key. - OpenRouter: Reported spend by API key, token detail by import. - Gemini CLI: Local history. - OpenCode: Local database. ## Included features - Usage and API-equivalent cost for six AI coding tools - Cache savings or overhead, net of write premiums - Claude Code and Codex limits with reset countdowns - All 17 native desktop widgets - Several accounts per tool - Monthly budget, any date range, CSV export - A license for one Mac, which you can move any time ## Frequently asked questions ### What is everyaitoken? everyaitoken is a native macOS app that brings AI coding usage, supported account limits, API-equivalent cost estimates, and cache savings together for Codex, Claude Code, Cursor, OpenRouter, Gemini CLI, and OpenCode. It includes a menu bar view, 17 desktop widgets, multiple accounts, and CSV export. It is not an AI model or a chat service. ### Does everyaitoken collect analytics? Optional product analytics use PostHog only in a configured build and only after you enable them in Settings. They contain basic app interactions, not usage history, prompts, code, credentials, account names, token counts, or costs. Website analytics are on by default in a configured build and can be stopped through Analytics settings in the footer. Both exclude session recordings and automatic interaction capture. ### How is this different from free trackers? Free trackers are good at counting tokens and showing limits. everyaitoken does that for six tools in one native app, then goes further on cost. It prices cache reads, 5-minute writes, and 1-hour writes at each provider’s rates, and shows what caching saved you after the write premiums, or what it cost you when a cache went unused. Reported spend stays separate from estimates, all 17 widgets are included, and it’s $9 once. It also has limits: it tracks six tools rather than dozens, and it shows session and weekly limits for Claude Code and Codex only. ### Which tools does it support? Codex by OpenAI, Claude Code by Anthropic, Cursor, OpenRouter, Gemini CLI by Google, and OpenCode. OpenRouter covers whichever models you route through it. You can connect more than one account for the same tool, such as a personal and a work account. ### Where do the limits come from? From the tools themselves, not estimated from log files. Codex limits come from your installed Codex, which reads them from OpenAI using the ChatGPT sign-in you already use with Codex, and refresh on the schedule you choose, from every minute to every 15 minutes. Claude Code limits come from the rate limits Claude Code passes to its status line, through an optional bridge that keeps your existing status line. They update while you use Claude Code, so they can fall behind while it’s idle. Once a reading is more than 15 minutes old, everyaitoken shows when it was last observed. ### Does it read or upload my prompts? It never stores or uploads them. To count tokens, everyaitoken parses your tools’ local session logs, which also contain your conversations, and keeps only usage metadata such as token counts, models, timestamps, and project folders. Your prompts, responses, and code are discarded, and there is no everyaitoken server to send anything to. The license check with Polar, the store that sells everyaitoken, sends your license key, a hashed ID for your Mac, the Mac’s name, and the app and macOS versions, never usage data. API keys you add and your license key stay in the macOS Keychain. ### Does it need my password or browser cookies? No passwords and no browser cookies. Codex limits use Codex’s own sign-in. Cursor uses the sign-in Cursor already saved on your Mac, or an Enterprise admin API key, which reports the whole team’s usage for the last 30 days. OpenRouter uses an API key you paste in. ### How is cost calculated? Each request is priced at its provider’s published API rates: OpenAI for Codex, Anthropic for Claude, Google for Gemini, and OpenRouter’s own per-model prices. Cache reads and cache writes are priced separately from ordinary input, and costs are computed in exact decimals. The result is an API-equivalent cost. When OpenRouter or Cursor reports its own spend figure, such as OpenRouter key spending or Cursor’s included-plan usage this cycle, everyaitoken shows it separately and never mixes it in. ### What does “cache savings” mean? AI coding tools resend your context with every request. Providers bill that repeated context as discounted cached input instead of full-price input. Writing the cache can cost extra: Anthropic charges 1.25× input for 5-minute writes and 2× for 1-hour writes. Cache savings compare what the tokens cost with caching against what the same tokens would cost as ordinary input, after those write premiums. ### Why can savings be negative? If a tool writes to the cache and the cache expires before anything reads it, you pay the write premium for nothing. everyaitoken shows that as cache overhead instead of hiding it, so you can see where caching worked against you. ### I’m on a ChatGPT or Claude subscription. Is it still useful? Yes. You see your Codex and Claude Code limits with exact reset times, with an optional alert at 80%. You also see what the same usage would cost at API prices. That figure is an API-equivalent estimate, not your bill, and it shows how much value your plan delivers. ### Is Cursor or OpenRouter spend real or estimated? Both are shown, and never mixed. everyaitoken shows OpenRouter’s spending for your API key and Cursor’s usage against your plan this cycle as each provider reports them. Neither is an invoice. Token-level costs are API-equivalent estimates. If a price is unknown, those tokens are left out of the estimate and the figure is flagged as unpriced or partial. ### Does it track the OpenAI API platform or ChatGPT? No. For OpenAI, everyaitoken tracks Codex usage and limits. It doesn’t read the OpenAI API platform dashboard or ChatGPT chat usage. ### How do the desktop widgets work? They are native macOS widgets. Right-click your desktop, choose Edit Widgets, and search for everyaitoken. Widgets show the latest data the app collected, and macOS decides when they refresh, so keep everyaitoken running in your menu bar. ### What are the system requirements? macOS 14 Sonoma or later on an Apple silicon Mac. Intel Macs aren’t supported. ### Is it a subscription? No. everyaitoken is $9 once, with a lifetime license for one Mac. ### Is there a free trial? No, and there is no free version. everyaitoken asks for a license key the first time you open it. ### How many Macs can I use it on? One at a time, and you can move your license any time. On the old Mac, choose Deactivate This Mac in everyaitoken’s settings, then enter your license key on the new one. If the old Mac is gone, remove it in Polar’s customer portal instead. ### Where do I find my license key? On the confirmation page right after you buy, and any time later in Polar’s customer portal, where you sign in with a code sent to the email you used at checkout. ## Official website resources - [Homepage](https://everyaitoken.com/) - [Product reference](https://everyaitoken.com/product) - [Free cache savings calculator](https://everyaitoken.com/tools/cache-savings): editable token counts and API rates, including cache write costs; API-equivalent estimates, not subscription bills. - [Privacy disclosures](https://everyaitoken.com/#privacy) - [Model comparisons](https://everyaitoken.com/blog) - [AI-readable index](https://everyaitoken.com/llms.txt) - [Expanded reference bundle](https://everyaitoken.com/llms-full.txt) ## Dated model price references Prices below are USD per million tokens at the listed standard rates. Observation dates identify when source rates were recorded, not a guarantee that rates remain current. Consult the original source before making a purchase decision. Notes include pricing qualifications. ### Claude Fable 5.1 Reference: https://everyaitoken.com/blog/models/claude-fable-5-1 Maker: Anthropic. API model ID: claude-fable-5-1. Status: current. Price observed: 2026-09-26. Input: $10. Output: $50. Cached input: $0.25. Cache write: $12.50. One-hour cache write: $20. Pricing source: [Anthropic: Pricing](https://platform.claude.com/docs/en/about-claude/pricing) - The full 1M context window is billed at standard rates. ### Claude Mythos 5.1 Reference: https://everyaitoken.com/blog/models/claude-mythos-5-1 Maker: Anthropic. API model ID: claude-mythos-5-1. Status: restricted. Price observed: 2026-09-26. Input: $10. Output: $50. Cached input: $0.25. Cache write: $12.50. One-hour cache write: $20. Pricing source: [Anthropic: Pricing](https://platform.claude.com/docs/en/about-claude/pricing) - The full 1M context window is billed at standard rates. ### Claude Opus 5.5 Reference: https://everyaitoken.com/blog/models/claude-opus-5-5 Maker: Anthropic. API model ID: claude-opus-5-5. Status: current. Price observed: 2026-09-26. Input: $4. Output: $20. Cached input: $0.20. Cache write: $5. One-hour cache write: $8. Pricing source: [Anthropic: Pricing](https://platform.claude.com/docs/en/about-claude/pricing) - The full 1M context window is billed at standard rates. - Fast mode, a research preview on the Claude API, costs $8 input and $40 output per million tokens. ### Claude Sonnet 5 Reference: https://everyaitoken.com/blog/models/claude-sonnet-5 Maker: Anthropic. API model ID: claude-sonnet-5. Status: current. Price observed: 2026-09-26. Input: $2. Output: $10. Cached input: $0.20. Cache write: $2.50. One-hour cache write: $4. Pricing source: [Anthropic: Pricing](https://platform.claude.com/docs/en/about-claude/pricing) - The full 1M context window is billed at standard rates. - The launch price became the standard price on August 10, 2026, and a planned increase was cancelled. ### Claude Haiku 4.5 Reference: https://everyaitoken.com/blog/models/claude-haiku-4-5 Maker: Anthropic. API model ID: claude-haiku-4-5. Status: current. Price observed: 2026-09-26. Input: $1. Output: $5. Cached input: $0.10. Cache write: $1.25. One-hour cache write: $2. Pricing source: [Anthropic: Pricing](https://platform.claude.com/docs/en/about-claude/pricing) ### Claude Fable 5 Reference: https://everyaitoken.com/blog/models/claude-fable-5 Maker: Anthropic. API model ID: claude-fable-5. Status: previous. Price observed: 2026-09-26. Input: $10. Output: $50. Cached input: $1. Cache write: $12.50. One-hour cache write: $20. Pricing source: [Anthropic: Pricing](https://platform.claude.com/docs/en/about-claude/pricing) - The full 1M context window is billed at standard rates. ### Claude Opus 5 Reference: https://everyaitoken.com/blog/models/claude-opus-5 Maker: Anthropic. API model ID: claude-opus-5. Status: previous. Price observed: 2026-09-26. Input: $5. Output: $25. Cached input: $0.50. Cache write: $6.25. One-hour cache write: $10. Pricing source: [Anthropic: Pricing](https://platform.claude.com/docs/en/about-claude/pricing) - The full 1M context window is billed at standard rates. - Fast mode, a research preview on the Claude API, costs $10 input and $50 output per million tokens. ### Claude Opus 4.8 Reference: https://everyaitoken.com/blog/models/claude-opus-4-8 Maker: Anthropic. API model ID: claude-opus-4-8. Status: previous. Price observed: 2026-09-26. Input: $5. Output: $25. Cached input: $0.50. Cache write: $6.25. One-hour cache write: $10. Pricing source: [Anthropic: Pricing](https://platform.claude.com/docs/en/about-claude/pricing) - The full 1M context window is billed at standard rates. - Fast mode, a research preview on the Claude API, costs $10 input and $50 output per million tokens. ### Claude Sonnet 4.6 Reference: https://everyaitoken.com/blog/models/claude-sonnet-4-6 Maker: Anthropic. API model ID: claude-sonnet-4-6. Status: previous. Price observed: 2026-09-26. Input: $3. Output: $15. Cached input: $0.30. Cache write: $3.75. One-hour cache write: $6. Pricing source: [Anthropic: Pricing](https://platform.claude.com/docs/en/about-claude/pricing) - The full 1M context window is billed at standard rates on the API. ### GPT-6 Astra Reference: https://everyaitoken.com/blog/models/gpt-6-astra Maker: OpenAI. API model ID: gpt-6-astra. Status: current. Price observed: 2026-09-28. Input: $10. Output: $50. Cached input: $1. Cache write: $12.50. Pricing source: [OpenAI: API pricing](https://developers.openai.com/api/docs/pricing) - Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. ### GPT-6 Sol Reference: https://everyaitoken.com/blog/models/gpt-6-sol Maker: OpenAI. API model ID: gpt-6-sol. Status: current. Price observed: 2026-09-28. Input: $2. Output: $10. Cached input: $0.20. Cache write: $2.50. Pricing source: [OpenAI: API pricing](https://developers.openai.com/api/docs/pricing) - Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. ### GPT-6 Luna Reference: https://everyaitoken.com/blog/models/gpt-6-luna Maker: OpenAI. API model ID: gpt-6-luna. Status: current. Price observed: 2026-09-28. Input: $0.10. Output: $0.50. Cached input: $0.01. Cache write: $0.125. Pricing source: [OpenAI: API pricing](https://developers.openai.com/api/docs/pricing) - Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. ### GPT-5.6 Sol Reference: https://everyaitoken.com/blog/models/gpt-5-6-sol Maker: OpenAI. API model ID: gpt-5.6-sol. Status: previous. Price observed: 2026-09-28. Input: $4. Output: $20. Cached input: $0.40. Cache write: $5. Pricing source: [OpenAI: API pricing](https://developers.openai.com/api/docs/pricing) - Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. - These are promotional rates, available at least through November 21, 2026. ### GPT-5.6 Terra Reference: https://everyaitoken.com/blog/models/gpt-5-6-terra Maker: OpenAI. API model ID: gpt-5.6-terra. Status: previous. Price observed: 2026-09-28. Input: $2. Output: $12. Cached input: $0.20. Cache write: $2.50. Pricing source: [OpenAI: API pricing](https://developers.openai.com/api/docs/pricing) - Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. ### GPT-5.6 Luna Reference: https://everyaitoken.com/blog/models/gpt-5-6-luna Maker: OpenAI. API model ID: gpt-5.6-luna. Status: previous. Price observed: 2026-09-28. Input: $0.20. Output: $1.20. Cached input: $0.02. Cache write: $0.25. Pricing source: [OpenAI: API pricing](https://developers.openai.com/api/docs/pricing) - Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. ### GPT-5.5 Reference: https://everyaitoken.com/blog/models/gpt-5-5 Maker: OpenAI. API model ID: gpt-5.5. Status: previous. Price observed: 2026-09-28. Input: $5. Output: $30. Cached input: $0.50. Cache write: $5. Pricing source: [OpenAI: API pricing](https://developers.openai.com/api/docs/pricing) - Prompts over 272K input tokens cost 2x for input and 1.5x for output, for the full session. ### GPT-5.4 Reference: https://everyaitoken.com/blog/models/gpt-5-4 Maker: OpenAI. API model ID: gpt-5.4. Status: previous. Price observed: 2026-09-28. Input: $2.50. Output: $15. Cached input: $0.25. Cache write: $2.50. Pricing source: [OpenAI: API pricing](https://developers.openai.com/api/docs/pricing) - Prompts over 272K input tokens cost 2x for input and 1.5x for output, for the full session. ### GPT-5.3-Codex Reference: https://everyaitoken.com/blog/models/gpt-5-3-codex Maker: OpenAI. API model ID: gpt-5.3-codex. Status: previous. Price observed: 2026-09-28. Input: $1.75. Output: $14. Cached input: $0.175. Cache write: $1.75. Pricing source: [OpenAI: API pricing](https://developers.openai.com/api/docs/pricing) ### GPT-5.2-Codex Reference: https://everyaitoken.com/blog/models/gpt-5-2-codex Maker: OpenAI. API model ID: gpt-5.2-codex. Status: previous. Price observed: 2026-09-28. Input: $1.75. Output: $14. Cached input: $0.175. Cache write: $1.75. Pricing source: [OpenAI docs: GPT-5.2-Codex](https://developers.openai.com/api/docs/models/gpt-5.2-codex) ### Gemini 3.8 Flash Reference: https://everyaitoken.com/blog/models/gemini-3-8-flash Maker: Google. API model ID: gemini-3.8-flash. Status: current. Price observed: 2026-09-28. Input: $0.75. Output: $3.75. Cached input: $0.075. Cache write: $0.75. Pricing source: [Google: Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing) - Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens. ### Gemini 3.1 Pro Preview Reference: https://everyaitoken.com/blog/models/gemini-3-1-pro-preview Maker: Google. API model ID: gemini-3.1-pro-preview. Status: preview. Price observed: 2026-09-28. Input: $2. Output: $12. Cached input: $0.20. Cache write: $2. Pricing source: [Google: Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing) - Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens. ### Gemini 3.5 Flash-Lite Reference: https://everyaitoken.com/blog/models/gemini-3-5-flash-lite Maker: Google. API model ID: gemini-3.5-flash-lite. Status: current. Price observed: 2026-09-28. Input: $0.30. Output: $2.50. Cached input: $0.03. Cache write: $0.30. Pricing source: [Google: Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing) ### Gemini 3.7 Flash Reference: https://everyaitoken.com/blog/models/gemini-3-7-flash Maker: Google. API model ID: gemini-3.7-flash. Status: previous. Price observed: 2026-09-28. Input: $0.75. Output: $3.75. Cached input: $0.075. Cache write: $0.75. Pricing source: [Google: Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing) - Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens. ### Gemini 3.6 Flash Reference: https://everyaitoken.com/blog/models/gemini-3-6-flash Maker: Google. API model ID: gemini-3.6-flash. Status: previous. Price observed: 2026-09-28. Input: $0.75. Output: $3.75. Cached input: $0.075. Cache write: $0.75. Pricing source: [Google: Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing) - Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens. ### Gemini 3.5 Flash Reference: https://everyaitoken.com/blog/models/gemini-3-5-flash Maker: Google. API model ID: gemini-3.5-flash. Status: previous. Price observed: 2026-09-28. Input: $1.50. Output: $9. Cached input: $0.15. Cache write: $1.50. Pricing source: [Google: Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing) ### Gemini 3.1 Flash-Lite Reference: https://everyaitoken.com/blog/models/gemini-3-1-flash-lite Maker: Google. API model ID: gemini-3.1-flash-lite. Status: previous. Price observed: 2026-09-28. Input: $0.25. Output: $1.50. Cached input: $0.025. Cache write: $0.25. Pricing source: [Google: Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing) ### Gemini 2.5 Pro Reference: https://everyaitoken.com/blog/models/gemini-2-5-pro Maker: Google. API model ID: gemini-2.5-pro. Status: previous. Price observed: 2026-09-28. Input: $1.25. Output: $10. Cached input: $0.125. Cache write: $1.25. Pricing source: [Google: Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing) - Prompts over 200K input tokens cost $2.50 input, $0.25 cached, and $15 output per million tokens. ### Gemini 2.5 Flash Reference: https://everyaitoken.com/blog/models/gemini-2-5-flash Maker: Google. API model ID: gemini-2.5-flash. Status: previous. Price observed: 2026-09-28. Input: $0.30. Output: $2.50. Cached input: $0.03. Cache write: $0.30. Pricing source: [Google: Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing) ### DeepSeek-V4.1-Flash Reference: https://everyaitoken.com/blog/models/deepseek-v4-1-flash Maker: DeepSeek. API model ID: deepseek-flash. Status: current. Price observed: 2026-09-28. Input: $0.30. Output: $1.20. Cached input: $0.006. Cache write: $0.30. Pricing source: [DeepSeek API: Models and pricing](https://api-docs.deepseek.com/quick_start/pricing) - Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. ### DeepSeek-V4-Pro Reference: https://everyaitoken.com/blog/models/deepseek-v4-pro Maker: DeepSeek. API model ID: deepseek-v4-pro. Status: current. Price observed: 2026-09-28. Input: $1.32. Output: $3.96. Cached input: $0.044. Cache write: $1.32. Pricing source: [DeepSeek API: Models and pricing](https://api-docs.deepseek.com/quick_start/pricing) - Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. ### GLM-5.3 Reference: https://everyaitoken.com/blog/models/glm-5-3 Maker: Z.ai. API model ID: glm-5.3. Status: current. Price observed: 2026-09-28. Input: $1.40. Output: $4.40. Cached input: $0.26. Cache write: $1.40. Pricing source: [Z.ai docs: Pricing](https://docs.z.ai/guides/overview/pricing) ### GLM-5.3-Flash Reference: https://everyaitoken.com/blog/models/glm-5-3-flash Maker: Z.ai. API model ID: glm-5.3-flash. Status: current. Price observed: 2026-09-28. Input: $0.15. Output: $0.50. Cached input: $0.03. Cache write: $0.15. Pricing source: [Z.ai docs: Pricing](https://docs.z.ai/guides/overview/pricing) ### Kimi K3 Reference: https://everyaitoken.com/blog/models/kimi-k3 Maker: Moonshot AI. API model ID: kimi-k3. Status: current. Price observed: 2026-09-28. Input: $3. Output: $15. Cached input: $0.30. Cache write: $3. Pricing source: [Kimi API: Chat pricing](https://platform.kimi.ai/docs/pricing/chat) - A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default. ### MiniMax M3 Reference: https://everyaitoken.com/blog/models/minimax-m3 Maker: MiniMax. API model ID: MiniMax-M3. Status: current. Price observed: 2026-09-28. Input: $0.30. Output: $1.20. Cached input: $0.06. Cache write: $0.30. Pricing source: [MiniMax docs: Pay-as-you-go pricing](https://platform.minimax.io/docs/guides/pricing-paygo) - MiniMax labels these rates a permanent 50% discount. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million tokens. ### Grok 4.7 Reference: https://everyaitoken.com/blog/models/grok-4-7 Maker: xAI. API model ID: grok-4.7. Status: current. Price observed: 2026-09-28. Input: $2. Output: $6. Cached input: $0.50. Cache write: $2. Pricing source: [xAI docs: Grok 4.7](https://docs.x.ai/developers/models/grok-4.7) - Once a prompt reaches 200K tokens, every token in the request costs $4 input, $1 cached, and $12 output per million. The US regional endpoint costs 10% more. ### Grok Build 0.1 Reference: https://everyaitoken.com/blog/models/grok-build-0-1 Maker: xAI. API model ID: grok-build-0.1. Status: preview. Price observed: 2026-09-28. Input: $1. Output: $2. Cached input: $0.20. Cache write: $1. Pricing source: [xAI docs: Grok Build 0.1](https://docs.x.ai/developers/models/grok-build-0.1) - Once a prompt reaches 200K tokens, every token in the request costs $2 input, $0.40 cached, and $4 output per million. ### Qwen3.8-Max Reference: https://everyaitoken.com/blog/models/qwen3-8-max Maker: Alibaba Qwen. API model ID: qwen3.8-max. Status: current. Price observed: 2026-09-28. Input: $2. Output: $6. Cached input: $0.25. Cache write: $2. Pricing source: [Qwen Cloud: Qwen3.8-Max](https://www.qwencloud.com/models/qwen3.8-max) - Qwen Cloud prices. Alibaba Cloud Model Studio's international scope lists the same input and output rates. ### Mistral Medium 3.5 Reference: https://everyaitoken.com/blog/models/mistral-medium-3-5 Maker: Mistral AI. API model ID: mistral-medium-3.5. Status: preview. Price observed: 2026-09-28. Input: $1.50. Output: $7.50. Cached input: $0.15. Cache write: $1.50. Pricing source: [Mistral docs: Mistral Medium 3.5](https://docs.mistral.ai/models/mistral-medium-3-5-26-04) - Mistral lists no per-model cache price. Its caching docs bill cached tokens at 10% of input, which is the $0.15 shown. ## Comparison answer summaries These are editorial summaries. Full articles contain context, scenarios, qualifications, and sources. ### Claude Fable 5 to Claude Fable 5.1: what the upgrade saves Source page: https://everyaitoken.com/blog/claude-fable-5-1-vs-claude-fable-5 Published: 2026-09-28. Updated: 2026-09-28. Claude Fable 5.1 costs the same as Claude Fable 5 for input, output, and cache writes, but a cache hit drops from $1 to $0.25 per million tokens. That makes the example agentic coding session $10.50 instead of $12.00, 13% less, while uncached work costs the same on both. Anthropic recommends moving to Fable 5.1, and Fable 5 stays available as a legacy model. ### Claude Fable 5.1 vs Claude Mythos 5.1: who can use which Source page: https://everyaitoken.com/blog/claude-fable-5-1-vs-claude-mythos-5-1 Published: 2026-09-28. Updated: 2026-09-28. Claude Mythos 5.1 is the same model as Claude Fable 5.1, with more permissive safeguards and access by invitation only. Their prices are identical, so the example agentic coding session costs $10.50 on either. Choose Mythos 5.1 only if your organization has been approved for it and your work needs those safeguards; everyone else uses Fable 5.1. ### Claude Fable 5.1 vs Claude Opus 5.5: when to pay 2.5x Source page: https://everyaitoken.com/blog/claude-fable-5-1-vs-claude-opus-5-5 Published: 2026-09-28. Updated: 2026-09-28. Claude Opus 5.5 is Anthropic's recommended starting model and the default in Claude Code, and Anthropic says it performs at the level of Claude Fable 5.1 on most work. Fable 5.1 lists at 2.5x the Opus 5.5 rates, and the example agentic coding session costs $10.50 on it against $4.40, a 2.4x gap. Anthropic suggests Fable 5.1 for the cases where Opus-tier results fall short. ### Claude Fable 5.1 vs Claude Sonnet 5 pricing, explained Source page: https://everyaitoken.com/blog/claude-fable-5-1-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, and the example agentic coding session costs $10.50 on it against $2.40, a 4.4x gap. Sonnet 5 is Anthropic's everyday balance of speed and intelligence, while Fable 5.1 is its most capable generally available model, which Anthropic suggests for work where Opus-tier results fall short. ### Claude Fable 5.1 vs Gemini 3.1 Pro Preview: 5.3x per session Source page: https://everyaitoken.com/blog/claude-fable-5-1-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.1 Pro Preview is the much cheaper model: the example agentic coding session costs $2.00 on it and $10.50 on Claude Fable 5.1, a 5.3x gap. Most of the difference is cache writes, which Google bills as ordinary input while Anthropic adds a premium. Fable 5.1 suits work that needs long outputs, flat long-context pricing, or a generally available model, and Gemini 3.1 Pro Preview suits Gemini CLI users who can accept a preview. ### Claude Fable 5.1 vs Gemini 3.8 Flash: the 14.8x question Source page: https://everyaitoken.com/blog/claude-fable-5-1-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash costs a small fraction of Claude Fable 5.1: the example agentic coding session is $0.71 on Flash and $10.50 on Fable 5.1, a 14.8x gap. Google pitches Flash for long-horizon software engineering at Flash prices, while Anthropic suggests Fable 5.1 for demanding work where Opus-tier results fall short. Flash fits high-volume agent loops and Fable 5.1 the hardest tasks, with one caveat: Flash's introductory rates end on December 31, 2026. ### Claude Fable 5.1 vs GPT-5.6 Sol: 2.5x apart, for now Source page: https://everyaitoken.com/blog/claude-fable-5-1-vs-gpt-5-6-sol Published: 2026-09-28. Updated: 2026-09-28. Claude Fable 5.1 costs 2.5x as much as GPT-5.6 Sol on every list price, and the example agentic coding session keeps that ratio at $10.50 against $4.20. GPT-5.6 Sol is OpenAI's previous flagship, on promotional rates available at least through November 21, 2026, and Codex now suggests GPT-6 Sol in its place. Choose Fable 5.1 for Anthropic's most capable model in Claude Code, and GPT-5.6 Sol if you rely on it in Codex cloud or through the gpt-5.6 API id. ### Claude Fable 5.1 vs GPT-6 Astra: why both cost $10.50 Source page: https://everyaitoken.com/blog/claude-fable-5-1-vs-gpt-6-astra Published: 2026-09-28. Updated: 2026-09-28. Claude Fable 5.1 and GPT-6 Astra have identical list prices, and the example agentic coding session costs $10.50 on each. Fable 5.1 charges $0.25 per million for a cache hit against $1 on Astra, but its 1-hour cache writes cost $20 against Astra's $12.50, and in this session the two differences cancel out. Pick Fable 5.1 for Claude Code and for prompts beyond 272K input tokens, which it bills at standard rates, and GPT-6 Astra for Codex, whose CLI ships it as the default model. ### Claude Fable 5.1 vs GPT-6 Sol: a steady 5x price gap Source page: https://everyaitoken.com/blog/claude-fable-5-1-vs-gpt-6-sol Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Sol costs a fifth as much as Claude Fable 5.1 on every list price, and the example agentic coding session keeps the 5x ratio: $2.10 against $10.50. Fable 5.1 is Anthropic's most capable generally available model, which Anthropic suggests when Opus-tier results fall short, while GPT-6 Sol is OpenAI's mid-priced GPT-6 model and the Codex docs' pick for complex coding. Use Sol for everyday volume in Codex, and keep Fable 5.1 for work where cheaper models have fallen short. ### Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite for sub-agents Source page: https://everyaitoken.com/blog/claude-haiku-4-5-vs-gemini-3-5-flash-lite Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.5 Flash-Lite is the cheaper sub-agent model, at $0.34 against $1.20 on Claude Haiku 4.5 for the example agentic coding session, a 3.5x gap. The gap shrinks to 2x on output-heavy work, because Flash-Lite's $2.50 output rate is half of Haiku 4.5's $5 while its input rate is less than a third. Haiku 4.5 runs in Claude Code, Cursor, and GitHub Copilot, and Flash-Lite runs in Gemini CLI but not in Cursor or Copilot. ### Claude Haiku 4.5 vs Gemini 3.8 Flash: cost now and in 2027 Source page: https://everyaitoken.com/blog/claude-haiku-4-5-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. Through December 31, 2026, Gemini 3.8 Flash is the cheaper model: the example agentic coding session costs $0.71 against $1.20 on Claude Haiku 4.5, a 1.7x gap. From January 1, 2027, Google lists $1.50 input, $0.15 cached, and $7.50 output for Gemini 3.8 Flash, all above Haiku 4.5's rates. Gemini 3.8 Flash also has a 1.05M context window against 200K, while Haiku 4.5 remains Anthropic's small model for Claude Code sub-agents. ### Claude Haiku 4.5 vs GPT-5.6 Luna for low-cost coding work Source page: https://everyaitoken.com/blog/claude-haiku-4-5-vs-gpt-5-6-luna Published: 2026-09-28. Updated: 2026-09-28. GPT-5.6 Luna is the cheaper model: $0.20 input and $1.20 output per million against $1 and $5 for Claude Haiku 4.5, and a month of example agentic coding sessions costs $24.20 against $132.00. Both already have successors in view, since Codex suggests moving from GPT-5.6 Luna to GPT-6 Luna, and Anthropic announced Claude Haiku 5.5 on September 22, 2026. Haiku 4.5 fits Claude Code sub-agents, and GPT-5.6 Luna fits OpenAI pipelines that already call it. ### Claude Haiku 4.5 vs GPT-5.6 Terra: small model, middle price Source page: https://everyaitoken.com/blog/claude-haiku-4-5-vs-gpt-5-6-terra Published: 2026-09-28. Updated: 2026-09-28. Claude Haiku 4.5 is the cheaper model, at $1 input and $5 output per million against $2 and $12 for GPT-5.6 Terra, and the example agentic coding session costs $1.20 against $2.20. GPT-5.6 Terra is OpenAI's middle GPT-5.6 tier with a 1.05M context window and 128K of output, against Haiku 4.5's 200K and 64K. Pick Haiku 4.5 for low-cost work that fits its window, and Terra when a task needs the extra room. ### Claude Haiku 4.5 vs GPT-6 Luna: a 10x gap in token prices Source page: https://everyaitoken.com/blog/claude-haiku-4-5-vs-gpt-6-luna Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Luna costs a tenth of Claude Haiku 4.5 per token, and a month of example agentic coding sessions comes to $11.55 against $132.00. Luna also offers a 1.05M context window and 128K of output against Haiku 4.5's 200K and 64K. Haiku 4.5 still earns its place as a sub-agent model inside Claude Code, though Anthropic lists its retirement as not sooner than October 15, 2026, with Claude Haiku 5.5 announced. ### Claude Opus 4.8 to Claude Opus 5.5: the cost of staying Source page: https://everyaitoken.com/blog/claude-opus-5-5-vs-claude-opus-4-8 Published: 2026-09-28. Updated: 2026-09-28. Claude Opus 5.5 costs less than Claude Opus 4.8 on every rate, and the example agentic coding session comes to $4.40 against $6.00, or $176.00 less over a month of 110 sessions. Anthropic recommends Opus 5.5 as its starting model for most work, but still recommends Opus 4.8 for cybersecurity work that needs reduced guardrails. ### Claude Opus 4.8 vs Gemini 3.1 Pro Preview: cost and limits Source page: https://everyaitoken.com/blog/claude-opus-4-8-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.1 Pro Preview costs far less: the example agentic coding session is $2.00 against $6.00 on Claude Opus 4.8, a 3x gap, and a month of sessions $220.00 against $660.00. Claude Opus 4.8 is now a legacy model that Anthropic still recommends for cybersecurity work needing reduced guardrails, while Gemini 3.1 Pro Preview is Google's current Pro model, still in preview. The gap narrows on prompts over 200K input tokens, where Gemini's rates rise and Opus 4.8 keeps its standard rates across its 1M window. ### Claude Opus 4.8 vs GPT-5.5: which costs less to code with? Source page: https://everyaitoken.com/blog/claude-opus-4-8-vs-gpt-5-5 Published: 2026-09-28. Updated: 2026-09-28. It depends on the shape of the work. GPT-5.5 costs less on the example agentic coding session, $5.00 against $6.00, because OpenAI bills cache writes as ordinary input while Anthropic charges up to 2x, but Claude Opus 4.8 costs less on output-heavy work because its output is $25 per million against $30. Both are previous-generation models, and GPT-5.5 leaves ChatGPT and Codex sign-in on October 14, 2026. ### Claude Opus 5 vs Claude Opus 4.8: same price, what changed Source page: https://everyaitoken.com/blog/claude-opus-5-vs-claude-opus-4-8 Published: 2026-09-28. Updated: 2026-09-28. Claude Opus 5 and Claude Opus 4.8 have identical prices, so the example agentic coding session costs $6.00 on either, and the choice comes down to what Anthropic says each is for. Anthropic pitched Opus 5 as close to Claude Fable 5 at half the price, and still recommends Opus 4.8 for cybersecurity work that needs reduced guardrails. Both are legacy models, and Anthropic's current recommendation for everyday Opus work is Claude Opus 5.5. ### Claude Opus 5 vs GPT-5.6 Sol: coding costs compared Source page: https://everyaitoken.com/blog/claude-opus-5-vs-gpt-5-6-sol Published: 2026-09-28. Updated: 2026-09-28. GPT-5.6 Sol lists 20% below Claude Opus 5 on every rate, and the example agentic coding session costs $4.20 on it against $6.00, because OpenAI's cache writes cost 1.25x input while Anthropic's 1-hour writes cost 2x. Both are previous-generation models: Anthropic recommends Claude Opus 5.5 over Opus 5, and Codex suggests GPT-6 Sol over GPT-5.6 Sol, whose current rates are promotional. ### Claude Opus 5.5 vs Claude Haiku 4.5: cost and context limits Source page: https://everyaitoken.com/blog/claude-opus-5-5-vs-claude-haiku-4-5 Published: 2026-09-28. Updated: 2026-09-28. Claude Haiku 4.5 lists at a quarter of Claude Opus 5.5's rates, but Opus 5.5's discounted cache hits narrow the example agentic coding session to 3.7x: $4.40 against $1.20. Opus 5.5 is Anthropic's recommended starting model and Claude Code's default, with a 1M context window and 128K of output. Haiku 4.5, capped at 200K of context and 64K of output, is Anthropic's cheapest current model, aimed at latency-sensitive work and coding sub-agents. ### Claude Opus 5.5 vs Claude Opus 5: is it worth switching? Source page: https://everyaitoken.com/blog/claude-opus-5-5-vs-claude-opus-5 Published: 2026-09-28. Updated: 2026-09-28. Claude Opus 5.5 replaces Claude Opus 5 as Anthropic's recommended Opus and as Claude Code's opus default, at 20% lower list prices and a cache hit that costs $0.20 instead of $0.50. On the example agentic coding session that is $4.40 against $6.00, 27% less. Opus 5 stays available as a legacy model, but it costs more on every line of the rate card. ### Claude Opus 5.5 vs Claude Sonnet 5: cost and caching Source page: https://everyaitoken.com/blog/claude-opus-5-5-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. Claude Sonnet 5 costs half as much per token as Claude Opus 5.5, but cache hits cost $0.20 per million on both. On a cache-heavy agentic coding session that narrows the gap to 1.8x, $2.40 against $4.40. Opus 5.5 is Anthropic's pick for long, sprawling jobs, and Sonnet 5 is the cheaper everyday choice with the same 1M context window. ### Claude Opus 5.5 vs Gemini 3.1 Pro Preview: cost and limits Source page: https://everyaitoken.com/blog/claude-opus-5-5-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.1 Pro Preview is cheaper: the example agentic coding session costs $2.00 on it and $4.40 on Claude Opus 5.5, a 2.2x gap, even though both charge $0.20 per million for a cache hit. Opus 5.5 is Claude Code's default and Anthropic's recommended starting model, with 128K of output and no long-context surcharge. Gemini 3.1 Pro Preview fits Gemini CLI users who want Google's Pro model at a lower price and can work with a preview. ### Claude Opus 5.5 vs GPT-5.5: the newer model costs less Source page: https://everyaitoken.com/blog/claude-opus-5-5-vs-gpt-5-5 Published: 2026-09-28. Updated: 2026-09-28. Claude Opus 5.5 costs less than GPT-5.5 on every workload here, from $4.40 against $5.00 on the example agentic coding session to $1.72 against $2.55 on output-heavy generation. The session gap is the narrowest because Anthropic's 1-hour cache writes cost more than GPT-5.5's writes, which OpenAI bills as ordinary input. Opus 5.5 is Claude Code's default and suits new work, while GPT-5.5 suits API workflows already built on it, since it leaves ChatGPT and Codex sign-in on October 14, 2026. ### Claude Opus 5.5 vs GPT-5.6 Sol: same list price, $0.20 apart Source page: https://everyaitoken.com/blog/claude-opus-5-5-vs-gpt-5-6-sol Published: 2026-09-28. Updated: 2026-09-28. Claude Opus 5.5 and GPT-5.6 Sol have the same list prices, $4 input and $20 output per million tokens, so uncached work costs the same on both. On the example agentic coding session GPT-5.6 Sol is $0.20 cheaper, $4.20 against $4.40, because Anthropic's 1-hour cache writes cost more than Opus 5.5's cheaper cache hits save. Choose Opus 5.5 in Claude Code, especially on the 5-minute cache, and GPT-5.6 Sol only where you already use it, since it runs on promotional rates and Codex now suggests GPT-6 Sol. ### Claude Opus 5.5 vs GPT-6 Astra: two tool defaults, priced Source page: https://everyaitoken.com/blog/claude-opus-5-5-vs-gpt-6-astra Published: 2026-09-28. Updated: 2026-09-28. Claude Opus 5.5 costs less: each GPT-6 Astra list price is 2.5x the Opus 5.5 rate, and the example agentic coding session costs $4.40 against $10.50, a 2.4x gap. Each is its tool's default, Opus 5.5 in Claude Code and GPT-6 Astra in Codex CLI's bundled model list, so for many developers the choice follows the tool. Anthropic says Opus 5.5 performs at the level of Claude Fable 5.1 on most work, and OpenAI pitches Astra for the hardest end-to-end work. ### Claude Opus 5.5 vs GPT-6 Sol: why caching widens the gap Source page: https://everyaitoken.com/blog/claude-opus-5-5-vs-gpt-6-sol Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Sol is the cheaper model at half Claude Opus 5.5's input, output, and cache-write prices, and the example agentic coding session costs $2.10 on Sol against $4.40 on Opus 5.5, a 2.1x gap. Both charge $0.20 per million for a cache hit, yet the session gap grows past 2x because Anthropic's 1-hour cache writes carry a heavier premium. Opus 5.5 is the natural pick in Claude Code, where it is the default, and Sol in Codex, whose docs recommend it for complex coding. ### Claude Sonnet 4.6 or GPT-5.3-Codex? Context and caching Source page: https://everyaitoken.com/blog/claude-sonnet-4-6-vs-gpt-5-3-codex Published: 2026-09-28. Updated: 2026-09-28. GPT-5.3-Codex costs less on the example agentic coding session, $1.93 against $3.60 on Claude Sonnet 4.6, because its input is cheaper and OpenAI adds no charge for cache writes. Sonnet 4.6 accepts 1M tokens of context, while GPT-5.3-Codex takes up to 272K input tokens of its 400K window. Both are previous-generation models, and GPT-5.3-Codex is no longer selectable in Codex with ChatGPT sign-in. ### Claude Sonnet 4.6 vs Gemini 2.5 Pro: prices and limits Source page: https://everyaitoken.com/blog/claude-sonnet-4-6-vs-gemini-2-5-pro Published: 2026-09-28. Updated: 2026-09-28. Gemini 2.5 Pro is cheaper on every rate, and the example agentic coding session costs $1.38 on it against $3.60 on Claude Sonnet 4.6. Sonnet 4.6 writes up to 128K tokens of output to Gemini 2.5 Pro's 65.5K and bills its full 1M context at standard rates, while Gemini 2.5 Pro's rates rise above 200K input tokens. Both are previous-generation models, and since September 18, 2026, Google limits Gemini 2.5 Pro to accounts that used it before. ### Claude Sonnet 4.6 vs GPT-5.4: API cost for coding agents Source page: https://everyaitoken.com/blog/claude-sonnet-4-6-vs-gpt-5-4 Published: 2026-09-28. Updated: 2026-09-28. GPT-5.4 costs less on the example agentic coding session, $2.50 against $3.60 on Claude Sonnet 4.6, mostly because OpenAI bills cache writes as ordinary input. Output costs $15 per million on both, so output-heavy work is nearly a tie at $1.28 and $1.29. Both are previous-generation models: Anthropic recommends Claude Sonnet 5, and GPT-5.4 left Codex's ChatGPT sign-in on August 31, 2026. ### Claude Sonnet 5 or Claude Haiku 4.5 for coding sub-agents? Source page: https://everyaitoken.com/blog/claude-sonnet-5-vs-claude-haiku-4-5 Published: 2026-09-28. Updated: 2026-09-28. Claude Haiku 4.5 costs half as much as Claude Sonnet 5 on every rate, so the example agentic coding session is $1.20 against $2.40. Sonnet 5 has a 1M context window to Haiku 4.5's 200K and writes up to 128K tokens of output to its 64K. Haiku 4.5 suits latency-sensitive work and sub-agents with small contexts, with the caveat that Anthropic lists its retirement as not sooner than October 15, 2026. ### Claude Sonnet 5 or GLM-5.3-Flash? What a 15x gap means Source page: https://everyaitoken.com/blog/glm-5-3-flash-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3-Flash costs $0.16 for the example agentic coding session against $2.40 on Claude Sonnet 5, a 15x gap that reaches 21.5x on output-heavy work, where Z.ai charges $0.50 per million output tokens against $10. Claude Sonnet 5 is the choice if you work in Claude Code or want the model Anthropic calls its most agentic Sonnet yet. GLM-5.3-Flash suits high-volume or budget-bound work through OpenRouter or OpenCode, and its MIT-licensed weights can be self-hosted. ### Claude Sonnet 5 vs Claude Sonnet 4.6: price and tokenizer Source page: https://everyaitoken.com/blog/claude-sonnet-5-vs-claude-sonnet-4-6 Published: 2026-09-28. Updated: 2026-09-28. Claude Sonnet 5 lists at $2 input and $10 output per million tokens, 33% below Claude Sonnet 4.6, and Anthropic calls it a drop-in upgrade. The example agentic coding session costs $2.40 on Sonnet 5 against $3.60, though Sonnet 5's newer tokenizer counts about 1.0 to 1.35x as many tokens for the same text, which eats into that gap. Sonnet 4.6 is now a legacy model, and Anthropic recommends the move. ### Claude Sonnet 5 vs Gemini 3.1 Pro Preview for coding costs Source page: https://everyaitoken.com/blog/claude-sonnet-5-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.1 Pro Preview costs less on the example agentic coding session, $2.00 against $2.40 for Claude Sonnet 5, because Google bills cache writes as ordinary input. Sonnet 5 costs less on uncached and output-heavy work and can write 128K tokens in one response against 65.5K. Gemini 3.1 Pro Preview is still a preview, its rates rise on prompts over 200K input tokens, and GitHub Copilot has already dropped it. ### Claude Sonnet 5 vs Gemini 3.5 Flash: legacy Flash on price Source page: https://everyaitoken.com/blog/claude-sonnet-5-vs-gemini-3-5-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.5 Flash costs less, $1.50 against $2.40 for Claude Sonnet 5 on the example agentic coding session, but the gap nearly closes on output-heavy work because its $9 output rate sits close to Sonnet 5's $10. Google now calls Gemini 3.5 Flash its legacy Flash, and while newer Flash models are on introductory rates it costs more per token than they do. For new work the real comparison is Sonnet 5 against a newer Flash, and Gemini 3.5 Flash matters mostly where Gemini CLI still falls back to it. ### Claude Sonnet 5 vs Gemini 3.8 Flash: price, limits, caching Source page: https://everyaitoken.com/blog/claude-sonnet-5-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash is much cheaper: its $0.75 input and $3.75 output rates are 63% below Claude Sonnet 5's, and the example agentic coding session costs $0.71 against $2.40, a 3.4x gap. Those are introductory rates that run through December 31, 2026, and Sonnet 5 can write up to 128K tokens in one response against 65.5K. Gemini 3.8 Flash suits low-cost work in Gemini CLI, and Sonnet 5 suits Claude Code and jobs that need long outputs. ### Claude Sonnet 5 vs GPT-5.5: API costs and the Codex exit Source page: https://everyaitoken.com/blog/claude-sonnet-5-vs-gpt-5-5 Published: 2026-09-28. Updated: 2026-09-28. Claude Sonnet 5 costs less on every workload compared here: the example agentic coding session is $2.40 against $5.00 on GPT-5.5, and output-heavy generation is $0.86 against $2.55. GPT-5.5 is OpenAI's April 2026 flagship, now a previous generation that leaves ChatGPT and Codex sign-in on October 14, 2026, while staying in the API. It mainly suits API users whose prompts are already tuned for it. ### Claude Sonnet 5 vs GPT-5.6 Sol: price, promo, and successor Source page: https://everyaitoken.com/blog/claude-sonnet-5-vs-gpt-5-6-sol Published: 2026-09-28. Updated: 2026-09-28. Claude Sonnet 5 costs half as much per token as GPT-5.6 Sol, and the example agentic coding session comes to $2.40 against $4.20, a 1.8x gap. GPT-5.6 Sol is a previous-generation model on promotional rates available at least through November 21, 2026, and outside Codex cloud the Codex docs suggest GPT-6 Sol in its place. Stay on GPT-5.6 Sol mainly if Codex cloud or code that calls the gpt-5.6 API id depends on it. ### Claude Sonnet 5 vs GPT-5.6 Terra: the cheaper pick flips Source page: https://everyaitoken.com/blog/claude-sonnet-5-vs-gpt-5-6-terra Published: 2026-09-28. Updated: 2026-09-28. Claude Sonnet 5 and GPT-5.6 Terra cost almost the same, and which is cheaper depends on the work. GPT-5.6 Terra is $0.20 cheaper on the example agentic coding session, $2.20 against $2.40, while Sonnet 5 is $0.16 cheaper on output-heavy generation because it charges $10 per million output tokens to Terra's $12. Terra is a previous model that Codex suggests replacing with GPT-6 Sol, so for new work the choice mostly comes down to your tools. ### Claude Sonnet 5 vs GPT-6 Astra: cost of OpenAI's top model Source page: https://everyaitoken.com/blog/claude-sonnet-5-vs-gpt-6-astra Published: 2026-09-28. Updated: 2026-09-28. Claude Sonnet 5 is the far cheaper model: every GPT-6 Astra rate is 5x higher, from $10 against $2 for input to $50 against $10 for output. On the example agentic coding session the gap narrows to 4.4x, $10.50 against $2.40, because Anthropic's 1-hour cache writes cost 2x input. OpenAI reserves GPT-6 Astra for the hardest long-running work and makes it the default in Codex CLI's bundled model list, so pick it for that tier of job and Sonnet 5 for everyday agentic coding. ### Claude Sonnet 5 vs GPT-6 Sol: same prices, different caches Source page: https://everyaitoken.com/blog/claude-sonnet-5-vs-gpt-6-sol Published: 2026-09-28. Updated: 2026-09-28. Claude Sonnet 5 and GPT-6 Sol list identical per-token prices, $2 input and $10 output per million, so uncached work costs the same on both. On the example agentic coding session GPT-6 Sol comes to $2.10 against $2.40, a $0.30 gap that comes entirely from Anthropic's pricier 1-hour cache writes. The practical choice is the tool you work in: Sonnet 5 lives in Claude Code, and GPT-6 Sol is the model the Codex docs recommend for complex coding. ### DeepSeek-V4-Pro vs Claude Opus 5.5: where 4.6x comes from Source page: https://everyaitoken.com/blog/deepseek-v4-pro-vs-claude-opus-5-5 Published: 2026-09-28. Updated: 2026-09-28. DeepSeek-V4-Pro costs $0.95 for the example agentic coding session against $4.40 on Claude Opus 5.5, a 4.6x gap in which cache writes account for $2.07 and output, at $20 against $3.96 per million, for most of the rest. Claude Opus 5.5 is the default model in Claude Code and Anthropic's recommended starting model for most work. DeepSeek-V4-Pro suits cost-driven agent work through OpenRouter, OpenCode, or self-hosting, though DeepSeek says a V4.1 Pro model is on the way. ### DeepSeek-V4-Pro vs GLM-5.3: similar rates, a 5.9x cache gap Source page: https://everyaitoken.com/blog/deepseek-v4-pro-vs-glm-5-3 Published: 2026-09-28. Updated: 2026-09-28. DeepSeek-V4-Pro and GLM-5.3 list nearly the same input and output prices, but a DeepSeek cache hit costs $0.044 per million against $0.26, so the example agentic coding session costs $0.95 on DeepSeek-V4-Pro against $1.44 on GLM-5.3. For cached agentic work DeepSeek-V4-Pro is the cheaper pick, and off-peak it costs 50% less again; GLM-5.3 suits you if you want Z.ai's GLM Coding Plan or its Anthropic-format endpoint. ### DeepSeek-V4-Pro vs GPT-6 Sol: API costs for agent work Source page: https://everyaitoken.com/blog/deepseek-v4-pro-vs-gpt-6-sol Published: 2026-09-28. Updated: 2026-09-28. DeepSeek-V4-Pro costs $0.95 for the example agentic coding session against $2.10 on GPT-6 Sol, a 2.2x gap that comes mostly from cache pricing: DeepSeek lists no write fee and charges $0.044 per million for a hit, against $0.20 on Sol. GPT-6 Sol is the model the Codex docs recommend for complex coding, while DeepSeek says it adapted V4-Pro for Codex through native Responses API support. Pick GPT-6 Sol to stay on OpenAI's own stack, and DeepSeek-V4-Pro for lower API costs or open weights you can host. ### DeepSeek-V4.1-Flash and GPT-5.6 Luna tie on session cost Source page: https://everyaitoken.com/blog/deepseek-v4-1-flash-vs-gpt-5-6-luna Published: 2026-09-28. Updated: 2026-09-28. DeepSeek-V4.1-Flash and GPT-5.6 Luna cost the same $0.22 for the example agentic coding session, and 110 sessions a month differ by $0.22, because Luna's lower input price and DeepSeek's lower cache-hit price cancel out. GPT-5.6 Luna is a previous-generation model that Codex suggests replacing with GPT-6 Luna, so it suits teams already running it in Codex, Cursor, or GitHub Copilot. DeepSeek-V4.1-Flash is the current model of the two, with open weights and a 50% off-peak discount that tips the session its way. ### DeepSeek-V4.1-Flash or Gemini 3.8 Flash? Price and caching Source page: https://everyaitoken.com/blog/deepseek-v4-1-flash-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. DeepSeek-V4.1-Flash costs less on every line, and the example agentic coding session comes to $0.22 against $0.71 on Gemini 3.8 Flash, a 3.2x gap, wider than on uncached work because DeepSeek's cache hits cost $0.006 per million. Gemini 3.8 Flash is the pick if you work in Gemini CLI or want Google's free tier, and its prices here are introductory rates that rise on January 1, 2027. DeepSeek-V4.1-Flash suits cost-driven work through OpenRouter or OpenCode, or on your own hardware with its MIT-licensed weights. ### DeepSeek-V4.1-Flash vs Claude Haiku 4.5: 1M context vs 200K Source page: https://everyaitoken.com/blog/deepseek-v4-1-flash-vs-claude-haiku-4-5 Published: 2026-09-28. Updated: 2026-09-28. DeepSeek-V4.1-Flash costs $0.22 for the example agentic coding session against $1.20 on Claude Haiku 4.5, a 5.5x gap, and it accepts 1M tokens of context where Haiku 4.5 stops at 200K. Claude Haiku 4.5 still makes sense inside Claude Code, where Anthropic positions it for latency-sensitive work and coding sub-agents, but Anthropic lists its retirement as not sooner than October 15, 2026 and has announced Claude Haiku 5.5. DeepSeek-V4.1-Flash fits work routed through OpenRouter or OpenCode, or self-hosted under its MIT license. ### DeepSeek-V4.1-Flash vs Claude Sonnet 5: a 10.9x cost gap Source page: https://everyaitoken.com/blog/deepseek-v4-1-flash-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. DeepSeek-V4.1-Flash prices the example agentic coding session at $0.22 against $2.40 on Claude Sonnet 5, a 10.9x gap that is wider than the 6.7x gap on the large uncached review because DeepSeek's cache hits cost $0.006 per million and writing the cache carries no premium. Pick Claude Sonnet 5 if you work in Claude Code or want what Anthropic calls its most agentic Sonnet yet, at its published rates. Pick DeepSeek-V4.1-Flash if cost per session drives the choice and you can run it through OpenRouter, OpenCode, or your own hardware, since its weights are open under the MIT license. ### DeepSeek-V4.1-Flash vs DeepSeek-V4-Pro: is Pro worth 4.3x? Source page: https://everyaitoken.com/blog/deepseek-v4-1-flash-vs-deepseek-v4-pro Published: 2026-09-28. Updated: 2026-09-28. DeepSeek-V4.1-Flash costs $0.22 for the example agentic coding session against $0.95 on DeepSeek-V4-Pro, and DeepSeek itself says tests by several parties put V4.1 Flash ahead of V4-Pro on performance, cost, speed, and total runtime. DeepSeek-V4-Pro keeps a case for setups built on its Responses API support and Codex adaptation, or on its low, high, and max effort levels. DeepSeek says V4-Pro service continues, billed as today, until a V4.1 Pro model arrives. ### DeepSeek-V4.1-Flash vs GLM-5.3-Flash: two open Flash models Source page: https://everyaitoken.com/blog/deepseek-v4-1-flash-vs-glm-5-3-flash Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3-Flash has the lower list prices, and the example agentic coding session costs $0.16 on it against $0.22 on DeepSeek-V4.1-Flash, a gap held to 1.4x because GLM's cache hits cost 5x as much as DeepSeek's. Off-peak, DeepSeek-V4.1-Flash charges 50% less, which is enough to reverse the session result. Both are MIT-licensed open-weight models that run through OpenRouter and OpenCode, so the choice turns on when you work, how much you write, and how much you cache. ### DeepSeek-V4.1-Flash vs GPT-6 Luna: which costs less? Source page: https://everyaitoken.com/blog/deepseek-v4-1-flash-vs-gpt-6-luna Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Luna is the cheaper of the two at published rates, $0.11 against $0.22 for the example agentic coding session, because its $0.10 input and $0.50 output undercut DeepSeek-V4.1-Flash on every uncached token. DeepSeek-V4.1-Flash wins back ground on cache hits and costs 50% less off-peak, which brings the session close to even. Choose on where you work: GPT-6 Luna runs in Codex, while DeepSeek-V4.1-Flash is an open-weight model you reach through OpenRouter, OpenCode, or your own servers. ### DeepSeek-V4.1-Flash vs MiniMax M3: same rates, cheaper hits Source page: https://everyaitoken.com/blog/deepseek-v4-1-flash-vs-minimax-m3 Published: 2026-09-28. Updated: 2026-09-28. DeepSeek-V4.1-Flash and MiniMax M3 list the same $0.30 input and $1.20 output per million, so uncached work costs the same on both, but a DeepSeek cache hit costs $0.006 against $0.06 and the example agentic coding session comes to $0.22 against $0.33. MiniMax M3 suits work with video input and lists a higher output limit, 524.3K tokens, though MiniMax recommends up to 131,072 per request, while DeepSeek-V4.1-Flash suits cache-heavy agents and off-peak work at 50% less. Both are open-weight models you reach through OpenRouter or OpenCode. ### Gemini 2.5 Pro vs Gemini 2.5 Flash after the access limit Source page: https://everyaitoken.com/blog/gemini-2-5-pro-vs-gemini-2-5-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 2.5 Pro costs about 4x Gemini 2.5 Flash on every rate, so the example agentic coding session comes to $1.38 against $0.34. Since September 18, 2026, Google limits access to both to accounts that used them before, though neither is deprecated. If you still have access, Pro is the one Google pitches for complex reasoning in code and Gemini CLI's Pro fallback, and Flash is its price-performance option for low-latency, high-volume tasks. ### Gemini 3.1 Pro Preview vs Gemini 2.5 Pro: preview or stable? Source page: https://everyaitoken.com/blog/gemini-3-1-pro-preview-vs-gemini-2-5-pro Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.1 Pro Preview is Google's current Pro model and costs more than Gemini 2.5 Pro: the example agentic coding session is $2.00 against $1.38. Gemini 2.5 Pro is the stable previous generation, but since September 18, 2026, Google limits it to accounts that used it before. Between the two, new accounts can only choose the Pro Preview, while accounts already on 2.5 Pro can keep a stable model at lower rates. ### Gemini 3.5 Flash to Gemini 3.8 Flash: far cheaper until 2027 Source page: https://everyaitoken.com/blog/gemini-3-8-flash-vs-gemini-3-5-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash is the newer model and, on its introductory rates, the cheaper one: the example agentic coding session costs $0.71 against $1.50 on Gemini 3.5 Flash. From January 1, 2027, 3.8 Flash moves to $1.50 input, $0.15 cached, and $7.50 output per million tokens, matching 3.5 Flash on input and caching and staying below its $9 output rate. Google now calls 3.5 Flash its legacy Flash, though Gemini CLI still uses it for some Sign in with Google accounts. ### Gemini 3.7 Flash vs Gemini 3.6 Flash for coding agents Source page: https://everyaitoken.com/blog/gemini-3-7-flash-vs-gemini-3-6-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.7 Flash and Gemini 3.6 Flash have the same introductory rates, so a cache-heavy agentic coding session costs $0.71 on either. Google pitched 3.6 Flash as a token-efficient Flash for general agentic and everyday work and 3.7 Flash for complex coding and multi-step execution. Both are previous-generation, neither is in Gemini CLI or Cursor, and GitHub Copilot plans to retire 3.6 Flash on October 2, 2026, in favor of Gemini 3.8 Flash. ### Gemini 3.8 Flash or Claude Opus 5.5 for agentic coding? Source page: https://everyaitoken.com/blog/claude-opus-5-5-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash is far cheaper: the example agentic coding session costs $0.71 on it and $4.40 on Claude Opus 5.5, a 6.2x gap, at introductory Flash rates that run through December 31, 2026. Opus 5.5 is Anthropic's recommended starting model and Claude Code's default, while Google aims Gemini 3.8 Flash at long-horizon software engineering and agents at Flash prices. Choose Flash for high-volume agent loops in Gemini CLI, and Opus 5.5 for Claude Code work or outputs longer than 65.5K tokens. ### Gemini 3.8 Flash or Gemini 3.5 Flash-Lite for subagents? Source page: https://everyaitoken.com/blog/gemini-3-8-flash-vs-gemini-3-5-flash-lite Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash costs 2.5x Gemini 3.5 Flash-Lite for input and 1.5x for output, so the example agentic coding session comes to $0.71 against $0.34. Google recommends the two together for new projects, with 3.8 Flash for long-horizon coding and agent work and 3.5 Flash-Lite for high-volume and subagent tasks. Gemini 3.8 Flash's introductory rates end on December 31, 2026, while Flash-Lite's rates carry no promotional note. ### Gemini 3.8 Flash vs Gemini 3.1 Pro Preview in Gemini CLI Source page: https://everyaitoken.com/blog/gemini-3-8-flash-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash costs about a third as much as Gemini 3.1 Pro Preview: the example agentic coding session is $0.71 against $2.00, and output-heavy work is 3.2x apart. Gemini CLI's default auto model uses both, as its Flash and Pro halves. Flash is on introductory rates that double on January 1, 2027, and the Pro Preview is still in preview, with Gemini 3.5 Pro announced but not yet released. ### Gemini 3.8 Flash vs Gemini 3.7 Flash at identical rates Source page: https://everyaitoken.com/blog/gemini-3-8-flash-vs-gemini-3-7-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash and Gemini 3.7 Flash share the same introductory rates, so the example agentic coding session costs $0.71 on either. The difference is support: Gemini 3.8 Flash is part of Gemini CLI's default auto model and is listed in Cursor, while Gemini 3.7 Flash is in neither. Google calls 3.8 Flash its most intelligent Flash model and 3.7 Flash its previous generation, so at equal prices the case for staying on 3.7 is a workflow already built around it. ### GLM-5.3 or Claude Sonnet 5: where a 40% saving comes from Source page: https://everyaitoken.com/blog/glm-5-3-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3 costs 40% less than Claude Sonnet 5 on the example agentic coding session, $1.44 against $2.40, and 55% less on output-heavy work, although Sonnet 5 charges less for each cache hit. Choose Sonnet 5 if you want it in Claude Code, Cursor, or GitHub Copilot; choose GLM-5.3 for the lower rates and open weights, reached through OpenRouter or OpenCode. ### GLM-5.3 vs Claude Opus 5.5: a 3.1x gap on coding sessions Source page: https://everyaitoken.com/blog/glm-5-3-vs-claude-opus-5-5 Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3 is the cheaper model by a wide margin: the example agentic coding session costs $1.44 on it against $4.40 on Claude Opus 5.5, a 3.1x gap, and output-heavy work is 4.4x apart. Pick Opus 5.5 if you work in Claude Code, where it is the default, or want the model Anthropic recommends for long, sprawling jobs; pick GLM-5.3 if cost leads the decision and you want open weights you can reach through OpenRouter or OpenCode. ### GLM-5.3 vs Gemini 3.1 Pro Preview: output price decides it Source page: https://everyaitoken.com/blog/glm-5-3-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3 costs $1.44 on the example agentic coding session against $2.00 on Gemini 3.1 Pro Preview, 28% less, and the gap grows to 2.6x on output-heavy work because Gemini output costs $12 per million against $4.40. Pick Gemini 3.1 Pro Preview if you work in Gemini CLI, where it is the Pro half of the default auto model; pick GLM-5.3 for cheaper output, a 128K output limit, and open weights. ### GLM-5.3 vs GLM-5.3-Flash: is Z.ai's flagship worth 9x? Source page: https://everyaitoken.com/blog/glm-5-3-vs-glm-5-3-flash Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3-Flash costs $0.16 on the example agentic coding session against $1.44 on GLM-5.3, a 9x gap, with the same 1M context, 128K output limit, and caching rules. Choose GLM-5.3 for the work Z.ai aims its flagship at, complex software engineering and long-horizon agents; choose GLM-5.3-Flash for high-volume or visual coding work, where Z.ai pitches its multimodal Flash model. ### GLM-5.3 vs GPT-6 Sol: open weights or Codex's coding pick Source page: https://everyaitoken.com/blog/glm-5-3-vs-gpt-6-sol Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3 costs $1.44 on the example agentic coding session against $2.10 on GPT-6 Sol, 31% less, and the gap reaches 55% on output-heavy work. Pick GPT-6 Sol if you code in Codex, whose docs recommend it for complex coding; pick GLM-5.3 if you want lower rates and open weights, through OpenRouter or OpenCode. ### GLM-5.3-Flash vs Claude Haiku 4.5: the 10x output gap Source page: https://everyaitoken.com/blog/glm-5-3-flash-vs-claude-haiku-4-5 Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3-Flash costs $0.16 for the example agentic coding session against $1.20 on Claude Haiku 4.5, and the gap is widest on output-heavy work, 10.8x, since Z.ai charges $0.50 per million output tokens against $5. Claude Haiku 4.5 fits Claude Code users who want Anthropic's low-cost model for sub-agents, though Anthropic has announced Claude Haiku 5.5 and lists Haiku 4.5's retirement as not sooner than October 15, 2026. GLM-5.3-Flash suits high-volume work through OpenRouter or OpenCode, with 1M tokens of context and MIT-licensed weights. ### GLM-5.3-Flash vs Gemini 3.8 Flash: prices now and in 2027 Source page: https://everyaitoken.com/blog/glm-5-3-flash-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3-Flash costs less on every line, $0.16 against $0.71 for the example agentic coding session, while Gemini 3.8 Flash costs 8x as much on output-heavy work, where it charges $3.75 per million output tokens against $0.50. Gemini 3.8 Flash is the choice inside Gemini CLI, Cursor, or GitHub Copilot, and its free tier covers it, but its current prices are introductory rates that end on December 31, 2026. GLM-5.3-Flash suits cost-driven work through OpenRouter or OpenCode, with MIT-licensed weights and 128K of output per response. ### GLM-5.3-Flash vs GPT-6 Luna: cache hits decide the price Source page: https://everyaitoken.com/blog/glm-5-3-flash-vs-gpt-6-luna Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Luna is the cheaper of the two at list prices: the example agentic coding session costs $0.11 against $0.16 on GLM-5.3-Flash, and most of the $0.05 gap is cache reads, where Luna charges $0.01 per million against $0.03. Both charge $0.50 per million output tokens, so output-heavy work costs the same $0.04 on each. Pick GPT-6 Luna for Codex or GitHub Copilot, and GLM-5.3-Flash for open MIT-licensed weights, visual coding, or work through OpenRouter and OpenCode. ### GPT-5.3-Codex vs Gemini 3.1 Pro Preview: context vs cost Source page: https://everyaitoken.com/blog/gpt-5-3-codex-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. Price barely separates them: the example agentic coding session costs $1.93 on GPT-5.3-Codex and $2.00 on Gemini 3.1 Pro Preview, while output-heavy work is cheaper on Gemini at $1.02 against $1.17. Limits matter more, since GPT-5.3-Codex accepts at most 272K input tokens and writes up to 128K, while Gemini 3.1 Pro Preview takes a 1.05M window but writes at most 65.5K. Pick GPT-5.3-Codex for existing API-key Codex workflows and long outputs, and Gemini 3.1 Pro Preview for very large prompts in Gemini CLI. ### GPT-5.3-Codex vs GPT-5.2-Codex: same price, what differs Source page: https://everyaitoken.com/blog/gpt-5-3-codex-vs-gpt-5-2-codex Published: 2026-09-28. Updated: 2026-09-28. GPT-5.3-Codex and GPT-5.2-Codex have identical rates, so every workload costs the same: $1.93 for the example agentic coding session on either. The choice comes down to OpenAI's claims, where GPT-5.3-Codex adds stronger reasoning and runs 25% faster for Codex users, and to availability. Both have left Codex for ChatGPT sign-in, and GPT-5.2-Codex has also left GitHub Copilot. ### GPT-5.4 vs GPT-5.3-Codex: does the 272K input cap matter? Source page: https://everyaitoken.com/blog/gpt-5-4-vs-gpt-5-3-codex Published: 2026-09-28. Updated: 2026-09-28. GPT-5.3-Codex is the cheaper of the two, $1.93 against $2.50 for the example agentic coding session, because GPT-5.4 costs 43% more for input but only 7% more for output. The bigger difference is context: GPT-5.3-Codex accepts up to 272K input tokens, while GPT-5.4 has a 1.05M window with higher rates above 272K. OpenAI folded GPT-5.3-Codex's coding into GPT-5.4, and both have left Codex for ChatGPT sign-in. ### GPT-5.5 vs Gemini 3.1 Pro Preview: a flat 2.5x gap Source page: https://everyaitoken.com/blog/gpt-5-5-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.1 Pro Preview costs less on every line: GPT-5.5's input, output, and cache prices are all 2.5x higher, and the example agentic coding session costs $5.00 on GPT-5.5 against $2.00. GPT-5.5 is OpenAI's April 2026 flagship and leaves ChatGPT and Codex sign-in on October 14, 2026, while Gemini 3.1 Pro Preview is Google's current Pro model, still in preview. Keep GPT-5.5 for API workflows that need its 128K output limit, and choose Gemini 3.1 Pro Preview for the lower cost in Gemini CLI. ### GPT-5.5 vs GPT-5.4 on the API: price, caching, retirement Source page: https://everyaitoken.com/blog/gpt-5-5-vs-gpt-5-4 Published: 2026-09-28. Updated: 2026-09-28. GPT-5.5 costs exactly 2x GPT-5.4 on input, output, and cache hits, and neither charges extra to write the cache, so the example agentic coding session costs $5.00 against $2.50. GPT-5.4 left Codex for ChatGPT sign-in on August 31, 2026, and GPT-5.5 follows on October 14, 2026, while both stay in the API. OpenAI calls GPT-5.5 a new class of intelligence for coding and GPT-5.4 its more affordable model, so the choice is how much of your work needs the pricier one. ### GPT-5.6 Luna vs Gemini 3.5 Flash-Lite: July's budget models Source page: https://everyaitoken.com/blog/gpt-5-6-luna-vs-gemini-3-5-flash-lite Published: 2026-09-28. Updated: 2026-09-28. GPT-5.6 Luna is the cheaper of the two, at $0.22 against $0.34 for the example agentic coding session and $0.10 against $0.21 for output-heavy generation. Gemini 3.5 Flash-Lite is Google's current low-cost tier, while GPT-5.6 Luna is a previous model that Codex suggests replacing with GPT-6 Luna. The tool you use settles most cases: Luna runs in Codex and Cursor, Flash-Lite in Gemini CLI. ### GPT-5.6 Sol vs Gemini 3.1 Pro Preview: two models in flux Source page: https://everyaitoken.com/blog/gpt-5-6-sol-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.1 Pro Preview is cheaper: the example agentic coding session costs $2.00 on it and $4.20 on GPT-5.6 Sol, a 2.1x gap, and 110 sessions a month come to $220.00 against $462.00. Both are in transition, since OpenAI has succeeded GPT-5.6 Sol with GPT-6 Sol and keeps it on promotional rates, while Google has announced Gemini 3.5 Pro. Pick GPT-5.6 Sol for 128K outputs and OpenAI's frontend and tool-calling strengths, and Gemini 3.1 Pro Preview for lower cost in Gemini CLI. ### GPT-5.6 Sol vs GPT-5.5: new cache write fee, lower rates Source page: https://everyaitoken.com/blog/gpt-5-6-sol-vs-gpt-5-5 Published: 2026-09-28. Updated: 2026-09-28. GPT-5.6 Sol charges 20% less than GPT-5.5 for input and 33% less for output, but it bills cache writes at 1.25x input while GPT-5.5 bills them as ordinary input. On the example agentic coding session that narrows the saving to 16%, $4.20 against $5.00. GPT-5.5 leaves ChatGPT and Codex sign-in on October 14, 2026, and GPT-5.6 Sol runs on promotional rates at least through November 21, 2026. ### GPT-5.6 Sol vs GPT-5.6 Terra: flagship or middle tier? Source page: https://everyaitoken.com/blog/gpt-5-6-sol-vs-gpt-5-6-terra Published: 2026-09-28. Updated: 2026-09-28. GPT-5.6 Sol costs 2x GPT-5.6 Terra for input and caching but only 1.7x for output, so the example agentic coding session costs $4.20 against $2.20. OpenAI pitched Sol as the GPT-5.6 flagship for complex professional work and Terra as the tier that balances intelligence and cost. Both are previous-generation models now, and Codex suggests moving from either one to GPT-6 Sol. ### GPT-5.6 Terra vs Gemini 3.8 Flash: rates, caching, status Source page: https://everyaitoken.com/blog/gpt-5-6-terra-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash is the cheaper model, with the example agentic coding session at $0.71 against $2.20 on GPT-5.6 Terra and output-heavy generation at $0.32 against $1.02. Its introductory rates end December 31, 2026, but the listed rates from January 1, 2027 ($1.50 input, $0.15 cached, $7.50 output) still sit below Terra's $2, $0.20, and $12. GPT-5.6 Terra is a previous model that Codex suggests replacing with GPT-6 Sol, so it mainly suits existing OpenAI setups. ### GPT-5.6 Terra vs GPT-5.6 Luna: a flat 10x after July's cuts Source page: https://everyaitoken.com/blog/gpt-5-6-terra-vs-gpt-5-6-luna Published: 2026-09-28. Updated: 2026-09-28. GPT-5.6 Terra costs exactly 10x GPT-5.6 Luna on every rate, so the example agentic coding session comes to $2.20 against $0.22. OpenAI positions Terra as the tier that balances intelligence and cost and Luna as the one for cost-sensitive, high-volume work. Codex now suggests different successors for them: GPT-6 Sol for Terra and GPT-6 Luna for Luna. ### GPT-6 Astra vs Gemini 3.1 Pro Preview: top models priced Source page: https://everyaitoken.com/blog/gpt-6-astra-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.1 Pro Preview costs far less: the example agentic coding session is $2.00 on it and $10.50 on GPT-6 Astra, a 5.3x gap, with input prices 5x apart. GPT-6 Astra, which OpenAI calls its most capable model, is also its most expensive and Codex CLI's default, and writes up to 128K tokens per response. Gemini 3.1 Pro Preview is Google's current Pro model, much cheaper but still a preview with a 65.5K output cap, and the Pro half of Gemini CLI's default. ### GPT-6 Astra vs Gemini 3.8 Flash: top tier or Flash tier? Source page: https://everyaitoken.com/blog/gpt-6-astra-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash is the budget choice by a wide margin: the example agentic coding session costs $0.71 on it and $10.50 on GPT-6 Astra, a 14.8x gap. Astra fits the hardest end-to-end work in Codex, where it is the CLI's default, and Flash fits high-volume agent loops in Gemini CLI. Flash's rates are introductory through December 31, 2026 and double after that, so budget with both price sets. ### GPT-6 Astra vs GPT-5.5: why sessions cost 2.1x, not 2x Source page: https://everyaitoken.com/blog/gpt-6-astra-vs-gpt-5-5 Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Astra costs 2x GPT-5.5 for input and 1.7x for output, but 2.5x for cache writes, because OpenAI bills writes at 1.25x input from GPT-5.6 on. That pushes a cache-heavy agentic coding session to $10.50 on Astra against $5.00 on GPT-5.5, a 2.1x gap. Astra is Codex CLI's bundled default, while GPT-5.5 exits ChatGPT and Codex sign-in on October 14, 2026. ### GPT-6 Astra vs GPT-6 Luna: the 100x spread in one lineup Source page: https://everyaitoken.com/blog/gpt-6-astra-vs-gpt-6-luna Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Astra costs 100x as much as GPT-6 Luna per token, so one example agentic coding session comes to $10.50 on Astra and $0.11 on Luna. OpenAI built Astra for the hardest end-to-end work and Luna for focused, high-volume tasks, so they suit different jobs rather than the same one. For complex coding, the Codex docs recommend neither of them but GPT-6 Sol, the middle tier. ### GPT-6 Astra vs GPT-6 Sol: Codex default or a fifth the cost? Source page: https://everyaitoken.com/blog/gpt-6-astra-vs-gpt-6-sol Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Astra charges exactly 5x the GPT-6 Sol rate on input, output, cache writes, and cache hits, so the example agentic coding session costs $10.50 on Astra and $2.10 on Sol. OpenAI pitches Astra for the hardest end-to-end work and Sol for complex coding and agent workflows, and the Codex docs recommend Sol for complex coding. Astra is still the default in Codex CLI's bundled model list, so check which one your sessions actually run on. ### GPT-6 Luna or Gemini 3.8 Flash for high-volume coding? Source page: https://everyaitoken.com/blog/gpt-6-luna-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Luna costs far less: its rates are 7.5x below Gemini 3.8 Flash's, and a month of example agentic coding sessions comes to $11.55 against $78.38. The two sit in different tiers, since OpenAI pitches Luna for focused, high-volume work while Google pitches Gemini 3.8 Flash for long-horizon software engineering and autonomous agents. Gemini 3.8 Flash is on introductory rates through December 31, 2026, after which its listed prices double. ### GPT-6 Luna vs Gemini 3.5 Flash-Lite: budget model costs Source page: https://everyaitoken.com/blog/gpt-6-luna-vs-gemini-3-5-flash-lite Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Luna is cheaper on every rate, at $0.10 input and $0.50 output per million against $0.30 and $2.50 for Gemini 3.5 Flash-Lite, and a month of example agentic coding sessions costs $11.55 against $36.85. The gap is widest on output-heavy work, since Flash-Lite's output rate is 5x Luna's. Luna fits Codex and GitHub Copilot, and Flash-Lite fits Gemini CLI and high-throughput sub-agents. ### GPT-6 Sol vs Gemini 3.1 Pro Preview: a near price tie Source page: https://everyaitoken.com/blog/gpt-6-sol-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. On price these two are almost level: the example agentic coding session costs $2.10 on GPT-6 Sol and $2.00 on Gemini 3.1 Pro Preview, while output-heavy generation costs $0.86 on Sol and $1.02 on Gemini. Limits and status separate them more than cost does, since GPT-6 Sol is a current release that writes up to 128K tokens and Gemini 3.1 Pro Preview is a preview capped at 65.5K. Sol fits Codex users and Gemini 3.1 Pro Preview fits Gemini CLI users. ### GPT-6 Sol vs Gemini 3.8 Flash: 3x now, less in 2027 Source page: https://everyaitoken.com/blog/gpt-6-sol-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash is cheaper: the example agentic coding session costs $0.71 on it and $2.10 on GPT-6 Sol, a 3x gap, and 110 sessions a month come to $78.38 against $231.00. The gap shrinks on January 1, 2027, when Flash's introductory rates end and its prices double to $1.50 input and $7.50 output per million tokens. Choose Sol for Codex and outputs up to 128K tokens, and Flash for high-volume agent loops in Gemini CLI. ### GPT-6 Sol vs GPT-5.6 Sol: half the price, same cache rules Source page: https://everyaitoken.com/blog/gpt-6-sol-vs-gpt-5-6-sol Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Sol costs half as much as GPT-5.6 Sol on input, output, cache writes, and cache hits, so the example agentic coding session drops from $4.20 to $2.10. Both bill cache writes at 1.25x input, so the saving holds on cached and uncached work alike. Codex suggests moving from GPT-5.6 Sol to GPT-6 Sol, though Codex cloud chats on ChatGPT plans still run GPT-5.6 Sol. ### GPT-6 Sol vs GPT-5.6 Terra: same input price, cheaper output Source page: https://everyaitoken.com/blog/gpt-6-sol-vs-gpt-5-6-terra Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Sol costs the same as GPT-5.6 Terra for input, cache writes, and cache hits, and 17% less for output, so the example agentic coding session costs $2.10 against $2.20. Codex suggests moving from GPT-5.6 Terra to GPT-6 Sol, which the Codex docs recommend for complex coding. Terra stays in the API, so the main reason to stay is a pipeline you have already tuned to it. ### GPT-6 Sol vs GPT-6 Luna: when a twentieth of the price fits Source page: https://everyaitoken.com/blog/gpt-6-sol-vs-gpt-6-luna Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Sol costs 20x as much as GPT-6 Luna per token, so a cache-heavy agentic coding session comes to $2.10 on Sol and $0.11 on Luna. The Codex docs recommend Sol for complex coding and Luna for focused, repeatable tasks. Running Sol on the hard problems and Luna on narrow, high-volume work is the split OpenAI's own positioning points to. ### Grok 4.7 vs Claude Opus 5.5: cheaper tokens, smaller window Source page: https://everyaitoken.com/blog/grok-4-7-vs-claude-opus-5-5 Published: 2026-09-28. Updated: 2026-09-28. On the example agentic coding session Grok 4.7 costs $2.30 against $4.40 on Claude Opus 5.5, and output-heavy work costs 3.2x as much on Opus 5.5 because its output rate is $20 per million against $6. Opus 5.5 keeps a 1M context window at standard rates, while Grok 4.7 stops at 500K and doubles every rate once a prompt reaches 200K tokens. Grok 4.7 suits cost-sensitive work in Cursor or GitHub Copilot, and Opus 5.5 suits long, sprawling jobs in Claude Code, which Anthropic names as a particular strength. ### Grok 4.7 vs Claude Sonnet 5: a near tie on cached sessions Source page: https://everyaitoken.com/blog/grok-4-7-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. Grok 4.7 and Claude Sonnet 5 both list $2 per million input tokens, and the example agentic coding session costs nearly the same on each: $2.30 on Grok 4.7 and $2.40 on Sonnet 5. The split widens on output-heavy work, where Grok 4.7's $6 output rate against $10 makes it 37% cheaper, while Sonnet 5's $0.20 cache hits favor long sessions that reread their context. Sonnet 5 also brings a 1M window and Claude Code, and Grok 4.7 stops at 500K with higher rates once a prompt reaches 200K tokens. ### Grok 4.7 vs Gemini 3.1 Pro Preview: two 200K price lines Source page: https://everyaitoken.com/blog/grok-4-7-vs-gemini-3-1-pro-preview Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.1 Pro Preview costs less on the example agentic coding session, $2.00 against $2.30 on Grok 4.7, because both bill cache writes as input and Gemini's cache hits cost $0.20 per million against $0.50. Grok 4.7 costs less wherever output dominates, at $6 per million against $12, so the output-heavy generation runs $0.54 against $1.02. Both raise their rates at 200K tokens, but Gemini 3.1 Pro Preview is still a preview with a 1.05M window, while Grok 4.7 is a current model with a 500K window. ### Grok 4.7 vs GPT-6 Sol: which costs less depends on the cache Source page: https://everyaitoken.com/blog/grok-4-7-vs-gpt-6-sol Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Sol costs less on the example agentic coding session, $2.10 against $2.30 on Grok 4.7, because its cache hits cost $0.20 per million against $0.50. Grok 4.7 costs less on everything uncached: its output rate is $6 against $10, so the output-heavy generation runs $0.54 against $0.86. The windows differ as well, 1.05M on GPT-6 Sol and 500K on Grok 4.7, and each has its own long-context price line, at 272K for GPT-6 Sol and 200K for Grok 4.7. ### Grok 4.7 vs Grok Build 0.1: which xAI model to code with Source page: https://everyaitoken.com/blog/grok-4-7-vs-grok-build-0-1 Published: 2026-09-28. Updated: 2026-09-28. Grok Build 0.1 costs less than Grok 4.7 on every example workload, $1.00 against $2.30 for the agentic coding session and $0.19 against $0.54 for output-heavy work, with input at half the price and output at a third. Grok 4.7 is xAI's current top model for coding and knowledge work, with a 500K window and a place in Cursor and GitHub Copilot. Grok Build 0.1 is xAI's early-access agentic coding model, with a 256K window, reachable through OpenRouter and OpenCode. ### Grok 4.7 vs Kimi K3: $6 or $15 per million output tokens Source page: https://everyaitoken.com/blog/grok-4-7-vs-kimi-k3 Published: 2026-09-28. Updated: 2026-09-28. Grok 4.7 costs less than Kimi K3 on every example workload, but by very different margins: $2.30 against $2.85 on the cached agentic coding session, and $0.54 against $1.29 on output-heavy work, since output costs $6 per million against $15. Kimi K3's cheaper cache hits, $0.30 against $0.50, are what keep the session close. Kimi K3 adds open weights, a 1.05M context window, and up to 1.05M tokens of output, while Grok 4.7 stops at 500K and doubles its rates once a prompt reaches 200K tokens. ### Grok Build 0.1 vs Claude Sonnet 5: price, window, and status Source page: https://everyaitoken.com/blog/grok-build-0-1-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. Grok Build 0.1 costs less than Claude Sonnet 5 on every example workload: $1.00 against $2.40 for the agentic coding session, and $0.19 against $0.86 for output-heavy work, since its output rate is $2 per million against $10. Sonnet 5 is a current model with a 1M context window that runs in Claude Code, Cursor, and GitHub Copilot, while Grok Build 0.1 is in early access with a 256K window and runs through OpenRouter and OpenCode. Grok Build 0.1 fits cost-driven agentic coding with prompts under 200K tokens, and Sonnet 5 fits large contexts and teams that work in Anthropic's tools. ### Grok Build 0.1 vs Gemini 3.8 Flash: prices now and in 2027 Source page: https://everyaitoken.com/blog/grok-build-0-1-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash costs less on the example agentic coding session, $0.71 against $1.00 on Grok Build 0.1, mostly because its cache hits cost $0.075 per million against $0.20. Grok Build 0.1 costs less on output-heavy work, $0.19 against $0.32, since its output rate is $2 against $3.75. Flash's rates are introductory through December 31, 2026, and double from January 1, 2027, which would put every Flash rate except cached reads above Grok Build 0.1's. ### Grok Build 0.1 vs GPT-6 Luna: $1.00 or $0.11 a session Source page: https://everyaitoken.com/blog/grok-build-0-1-vs-gpt-6-luna Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Luna is far cheaper: the example agentic coding session costs $0.11 on it against $1.00 on Grok Build 0.1, and a month of 110 sessions costs $11.55 against $110.00. The gap is narrowest on output-heavy work, 4.8x, because Grok Build 0.1's $2 output rate is 4x Luna's $0.50, while its cache hits cost 20x as much. OpenAI pitches Luna for focused, high-volume, repeatable work, and xAI describes Grok Build 0.1 as a coding model trained for agentic workflows, still in early access with a 256K window. ### Is DeepSeek-V4-Pro cheaper than Claude Sonnet 5? Source page: https://everyaitoken.com/blog/deepseek-v4-pro-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. Yes: DeepSeek-V4-Pro costs $0.95 for the example agentic coding session against $2.40 on Claude Sonnet 5, a 2.5x gap, and most of the difference is cache writes, which DeepSeek bills as ordinary input. Claude Sonnet 5 is Anthropic's balance of speed and intelligence and runs in Claude Code, Cursor, and GitHub Copilot. DeepSeek-V4-Pro fits agent work through OpenRouter or OpenCode, or self-hosted under the MIT license, with the caveat that DeepSeek plans a V4.1 Pro successor. ### Kimi K3 or GPT-6 Sol: which costs less for agentic coding? Source page: https://everyaitoken.com/blog/kimi-k3-vs-gpt-6-sol Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Sol is cheaper on every rate, and the example agentic coding session costs $2.10 on it against $2.85 on Kimi K3, a 26% saving. Pick GPT-6 Sol if you work in Codex, whose docs recommend it for complex coding; pick Kimi K3 for open weights or for responses longer than 128K tokens. ### Kimi K3 vs Claude Opus 5.5: similar caches, different prices Source page: https://everyaitoken.com/blog/kimi-k3-vs-claude-opus-5-5 Published: 2026-09-28. Updated: 2026-09-28. Kimi K3 costs $2.85 on the example agentic coding session against $4.40 on Claude Opus 5.5, 35% less, even though its cache hits cost $0.30 per million against $0.20. Choose Opus 5.5 for Claude Code, where it is the default, and for Anthropic's long-running agentic work; choose Kimi K3 for lower list prices, open weights, and responses of up to 1.05M tokens. ### Kimi K3 vs Claude Sonnet 5: when the open model costs more Source page: https://everyaitoken.com/blog/kimi-k3-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. Claude Sonnet 5 is the cheaper model here: Kimi K3's input, output, and cache-hit rates are all 1.5x higher, and the example agentic coding session costs $2.85 on Kimi K3 against $2.40 on Sonnet 5. The session gap is only 19% because Kimi's 5-minute cache writes cost less than Anthropic's mix of 5-minute and 1-hour writes, so pick Kimi K3 for open weights, very long outputs, or Moonshot's flagship, and Sonnet 5 for price and Claude Code. ### Kimi K3 vs DeepSeek-V4-Pro: why DeepSeek costs a third Source page: https://everyaitoken.com/blog/kimi-k3-vs-deepseek-v4-pro Published: 2026-09-28. Updated: 2026-09-28. DeepSeek-V4-Pro costs a third as much as Kimi K3 on the example agentic coding session, $0.95 against $2.85, and its off-peak hours cut its rates by 50% again. Pick Kimi K3 if you want it inside Cursor or GitHub Copilot or need outputs beyond 384K tokens; pick DeepSeek-V4-Pro when cost leads, especially for cache-heavy agent loops. ### Kimi K3 vs GLM-5.3: open-weight flagships compared on cost Source page: https://everyaitoken.com/blog/kimi-k3-vs-glm-5-3 Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3 costs about half as much as Kimi K3 on the example agentic coding session, $1.44 against $2.85, while output-heavy work costs 3.3x as much on Kimi K3, although the two charge almost the same for a cache hit. Pick Kimi K3 if you want it in Cursor or GitHub Copilot or need responses beyond 128K tokens; pick GLM-5.3 if price leads and OpenRouter or OpenCode fit your workflow. ### Kimi K3 vs GPT-6 Astra: two flagships, a 3.7x cost gap Source page: https://everyaitoken.com/blog/kimi-k3-vs-gpt-6-astra Published: 2026-09-28. Updated: 2026-09-28. Kimi K3 is far cheaper: the example agentic coding session costs $2.85 on it against $10.50 on GPT-6 Astra, and Astra's input, output, and cache-hit rates are all 3.3x higher. Pick GPT-6 Astra if you work in Codex, where it is the default in the CLI's bundled model list, and you accept OpenAI's claim that it finishes tasks with fewer output tokens; pick Kimi K3 for cost, open weights, and outputs beyond 128K. ### MiniMax M3 vs Claude Sonnet 5: Sonnet costs 7.3x as much Source page: https://everyaitoken.com/blog/minimax-m3-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. MiniMax M3 costs about a seventh as much as Claude Sonnet 5 on the example agentic coding session, $0.33 against $2.40, and the gap stays between 6.7x and 7.8x on every workload. Both accept 1M tokens of context, but MiniMax M3 raises its rates above 512K input tokens and ships open weights, while Sonnet 5 runs in Claude Code, and in Cursor and GitHub Copilot, which don't list MiniMax M3. Choose MiniMax M3 when cost per token decides, and Sonnet 5 when you want Anthropic's model inside those tools. ### MiniMax M3 vs Gemini 3.8 Flash: permanent discount vs promo Source page: https://everyaitoken.com/blog/minimax-m3-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. MiniMax M3 costs less than Gemini 3.8 Flash on every example workload, $0.33 against $0.71 for the agentic coding session and $0.11 against $0.32 for output-heavy work. The gap is smallest on cache reads, $0.06 against $0.075 per million, and it would widen at Flash's 2027 rates, since Google's introductory prices end on December 31, 2026, while MiniMax labels its own rates a permanent 50% discount. Gemini 3.8 Flash counters with Gemini CLI, Cursor, GitHub Copilot, and a free API tier, and MiniMax M3 with open weights and up to 524.3K tokens of output. ### MiniMax M3 vs GLM-5.3-Flash: open-weight pricing compared Source page: https://everyaitoken.com/blog/minimax-m3-vs-glm-5-3-flash Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3-Flash is the cheaper of these two open-weight models on every line, and the example agentic coding session costs $0.16 on it against $0.33 on MiniMax M3. MiniMax M3 makes sense if you need outputs beyond 128K tokens, up to 524.3K, or want the video input and desktop computer use MiniMax lists among its strengths. Both makers price a cache hit at 20% of input and charge nothing to write the cache, so the 2.1x session gap tracks their list prices. ### MiniMax M3 vs GPT-6 Luna: cents per session, 3x apart Source page: https://everyaitoken.com/blog/minimax-m3-vs-gpt-6-luna Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Luna costs about a third as much as MiniMax M3 on every example workload: $0.11 against $0.33 for the agentic coding session, and $11.55 against $36.30 for a month of 110 sessions. The biggest single difference is the cache hit, $0.01 per million on Luna against $0.06 on MiniMax M3. MiniMax M3 offers open weights, up to 524.3K tokens of output, and a long-context line at 512K rather than Luna's 272K, while Luna runs in Codex and GitHub Copilot. ### Mistral Medium 3.5 vs Claude Sonnet 5: the cache-write gap Source page: https://everyaitoken.com/blog/mistral-medium-3-5-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. Mistral Medium 3.5 lists its input, output, and cache-hit rates 25% below Claude Sonnet 5, and the example agentic coding session costs $1.43 against $2.40, 40% less, because Mistral charges no cache-write fee while Anthropic bills writes at 1.25x and 2x input. Sonnet 5 is a current model with a 1M context window that runs in Claude Code, Cursor, and GitHub Copilot. Mistral Medium 3.5 is a public preview with a 256K window and open weights, reachable here through OpenRouter, and it powers Mistral's own Vibe coding agents. ### Mistral Medium 3.5 vs Gemini 3.8 Flash: the 2027 price match Source page: https://everyaitoken.com/blog/mistral-medium-3-5-vs-gemini-3-8-flash Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.8 Flash costs exactly half as much as Mistral Medium 3.5 on every rate today, so the example agentic coding session runs $0.71 against $1.43. That changes on January 1, 2027, when Flash's introductory rates end and its new list prices, $1.50 input, $0.15 cached, and $7.50 output, equal Mistral Medium 3.5's rates. From then on the choice rests on limits and tools: Flash has a 1.05M window and runs in Gemini CLI, Cursor, and GitHub Copilot, while Mistral Medium 3.5 has open weights and a 256K window. ### Mistral Medium 3.5 vs GPT-6 Sol: 32% less per session Source page: https://everyaitoken.com/blog/mistral-medium-3-5-vs-gpt-6-sol Published: 2026-09-28. Updated: 2026-09-28. Mistral Medium 3.5 costs $1.43 on the example agentic coding session against $2.10 on GPT-6 Sol, 32% less, from rates 25% lower on input, output, and cache hits plus the absence of a cache-write premium. GPT-6 Sol brings a 1.05M context window, Codex, and GitHub Copilot, and the Codex docs recommend it for complex coding. Mistral Medium 3.5 is a public preview with a 256K window and open weights, reachable here through OpenRouter and used in Mistral's Vibe agents. ### Qwen3.8-Max or Claude Sonnet 5? Same input price, compared Source page: https://everyaitoken.com/blog/qwen3-8-max-vs-claude-sonnet-5 Published: 2026-09-28. Updated: 2026-09-28. Qwen3.8-Max and Claude Sonnet 5 charge the same $2 per million input tokens, so the gap comes from output, $6 against $10, and from cache writes, and the example agentic coding session costs $1.80 on Qwen3.8-Max against $2.40 on Sonnet 5. Pick Sonnet 5 if you work in Claude Code, Cursor, or GitHub Copilot; pick Qwen3.8-Max for the lower session cost through OpenRouter or OpenCode. ### Qwen3.8-Max vs Claude Opus 5.5: closed models, 2.4x apart Source page: https://everyaitoken.com/blog/qwen3-8-max-vs-claude-opus-5-5 Published: 2026-09-28. Updated: 2026-09-28. Qwen3.8-Max costs $1.80 on the example agentic coding session against $4.40 on Claude Opus 5.5, 59% less, with most of the gap coming from Anthropic's cache-write premium and its $20 output rate. Pick Opus 5.5 if you work in Claude Code, Cursor, or GitHub Copilot; pick Qwen3.8-Max for the lower price through OpenRouter or OpenCode, knowing it is a closed API model just as Opus 5.5 is. ### Qwen3.8-Max vs GLM-5.3: list prices, caching, and weights Source page: https://everyaitoken.com/blog/qwen3-8-max-vs-glm-5-3 Published: 2026-09-28. Updated: 2026-09-28. GLM-5.3 costs less than Qwen3.8-Max on every example workload, $1.44 against $1.80 for the agentic coding session and $0.25 against $0.36 for a large one-off review, because its input and output rates are $1.40 and $4.40 per million against $2 and $6. Cache hits are nearly identical, $0.26 on GLM-5.3 and $0.25 on Qwen3.8-Max, so the session gap, 20%, is narrower than the rate gap. GLM-5.3 ships open weights under Z.ai's own license, while Qwen3.8-Max is a closed API model whose base variant has open weights. ### Qwen3.8-Max vs GPT-6 Sol: equal input, $6 or $10 output Source page: https://everyaitoken.com/blog/qwen3-8-max-vs-gpt-6-sol Published: 2026-09-28. Updated: 2026-09-28. Qwen3.8-Max and GPT-6 Sol both charge $2 per million input tokens, and the example agentic coding session costs $1.80 on Qwen3.8-Max against $2.10 on GPT-6 Sol, a 14% gap driven by output and OpenAI's 1.25x cache-write charge. Pick GPT-6 Sol if you work in Codex, whose docs recommend it for complex coding; pick Qwen3.8-Max for cheaper output through OpenRouter or OpenCode. ### Qwen3.8-Max vs Kimi K3: a closed API against open weights Source page: https://everyaitoken.com/blog/qwen3-8-max-vs-kimi-k3 Published: 2026-09-28. Updated: 2026-09-28. Qwen3.8-Max is cheaper on every rate, and the example agentic coding session costs $1.80 on it against $2.85 on Kimi K3, 37% less. Pick Kimi K3 if you want open weights, access in Cursor or GitHub Copilot, or responses longer than 131K tokens; pick Qwen3.8-Max if a hosted API is all you need and you want the lower price through OpenRouter, OpenCode, or Qwen Cloud. ### Replacing Gemini 3.1 Flash-Lite with Gemini 3.5 Flash-Lite Source page: https://everyaitoken.com/blog/gemini-3-5-flash-lite-vs-gemini-3-1-flash-lite Published: 2026-09-28. Updated: 2026-09-28. Gemini 3.5 Flash-Lite is Google's named replacement for Gemini 3.1 Flash-Lite, which shuts down on May 7, 2027. The newer model costs 20% more for input and 67% more for output, so the example agentic coding session rises from $0.25 to $0.34. Google positions 3.5 Flash-Lite for subagent tasks and high-throughput execution, while 3.1 Flash-Lite was built for lightweight, high-frequency jobs such as routing and extraction. ### Upgrading from GPT-5.6 Luna to GPT-6 Luna: what changes Source page: https://everyaitoken.com/blog/gpt-6-luna-vs-gpt-5-6-luna Published: 2026-09-28. Updated: 2026-09-28. GPT-6 Luna costs half as much as GPT-5.6 Luna for input and caching and 58% less for output, so the example agentic coding session drops from $0.22 to $0.11. Codex suggests the move, while GPT-5.6 Luna stays available in the API. At these prices the gap is $12.65 over 110 sessions, so it matters most in high-volume pipelines.