AI Model Pricing & Specs Tracker
A running comparison of pricing ($/1M tokens), context windows, and release dates for the API models offered by major AI labs — OpenAI, Anthropic, Google, xAI and more. Every row links to its source and the date it was last verified. Updated weekly.
Last updated: 2026-08-10
API Cost Calculator
Enter your expected monthly input and output tokens to see an estimated monthly cost for every model with published pricing, ranked from cheapest to most expensive. As a rule of thumb, 1 million tokens is roughly 750,000 English words.
| Model | Provider | Estimated monthly cost |
|---|---|---|
| GPT-5.6 Luna | OpenAI | $4.40 |
| Gemini 3.5 Flash-Lite | $8.00 | |
| Grok 4.3 | SpaceXAI | $17.50 |
| Claude Haiku 4.5 | Anthropic | $20.00 |
| Muse Spark 1.1 | Meta | $21.00 |
| Gemini 3.6 Flash | $30.00 | |
| Qwen3.8-Max | Alibaba | $32.00 |
| Grok 4.5 | SpaceXAI | $32.00 |
| Gemini 3.5 Flash | $33.00 | |
| Claude Sonnet 5 | Anthropic | $40.00 |
| GPT-5.6 Terra | OpenAI | $44.00 |
| Gemini 3.1 Pro Preview | $44.00 | |
| Kimi K3 | Moonshot AI | $60.00 |
| Grok 4.5 Fast | SpaceXAI | $76.00 |
| Claude Opus 5 | Anthropic | $100.00 |
| Claude Opus 4.8 | Anthropic | $100.00 |
| Claude Opus 4.7 | Anthropic | $100.00 |
| GPT-5.6 Sol | OpenAI | $110.00 |
| Claude Fable 5 | Anthropic | $200.00 |
| Claude Mythos 5 | Anthropic | $200.00 |
Excluded due to unpublished pricing: 3 model(s)
* Rough estimate only — does not account for prompt caching or batch discounts. See each provider's source link for exact pricing.
All models
| Model | Provider | Tier | Context | Max output | Input ($/1M tokens) | Output ($/1M tokens) | Released | Source |
|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | Flagship | — | — | $5.00 | $30.00 | 2026-07-09 | Source as of 2026-08-10 |
| Notes Flagship of the GPT-5.6 family; fast mode is billed at double the standard rate ($10 input / $60 output). Context window and max output are not listed on the official pricing page. | ||||||||
| GPT-5.6 Terra | OpenAI | Mid-tier | — | — | $2.00 | $12.00 | 2026-07-09 | Source as of 2026-08-10 |
| Notes Balanced everyday-work tier; price cut 20% on 2026-07-30. Fast mode is billed at double the standard rate. Context window and max output are not listed on the official pricing page. | ||||||||
| GPT-5.6 Luna | OpenAI | Small | — | — | $0.20 | $1.20 | 2026-07-09 | Source as of 2026-08-10 |
| Notes Most cost-efficient tier; price cut 80% on 2026-07-30. Fast mode is billed at double the standard rate. Context window and max output are not listed on the official pricing page. | ||||||||
| Claude Fable 5 | Anthropic | Flagship | 1M | 128K | $10.00 | $50.00 | 2026-06-09 | Source as of 2026-08-10 |
| Notes Anthropic's most capable widely released model, GA on 2026-06-09. It uses the tokenizer introduced with Opus 4.7, which produces roughly 30% more tokens for the same text — worth noting when comparing costs. Price and specs confirmed in the official docs. | ||||||||
| Claude Mythos 5 | Anthropic | Flagship | 1M | 128K | $10.00 | $50.00 | 2026-06-09 | Source as of 2026-08-10 |
| Notes Shares Fable 5's specs and pricing; offered under limited availability for defensive cybersecurity work through Project Glasswing, by invitation only with no self-serve sign-up. | ||||||||
| Claude Opus 5 | Anthropic | Flagship | 1M | 128K | $5.00 | $25.00 | 2026-07-24 | Source as of 2026-08-10 |
| Notes New-generation Opus at the same price as Opus 4.8. Fast mode (research preview) is $10 input / $50 output, double the standard rate. Note that the Opus 4.7-era tokenizer produces roughly 30% more tokens for the same text. | ||||||||
| Claude Opus 4.8 | Anthropic | Flagship | 1M | 128K | $5.00 | $25.00 | — | Source as of 2026-08-10 |
| Notes A key baseline model across vendor benchmarks, now listed under legacy models in the official docs. Fast mode (research preview) is $10 input / $50 output. Release date not verified here. | ||||||||
| Claude Opus 4.7 | Anthropic | Flagship | 1M | 128K | $5.00 | $25.00 | — | Source as of 2026-08-10 |
| Notes Prior-generation flagship, listed under legacy models in the official docs. It introduced the newer tokenizer that produces roughly 30% more tokens for the same text. Fast mode is not supported (requests with speed: "fast" return an error). Release date not verified here. | ||||||||
| Claude Sonnet 5 | Anthropic | Mid-tier | 1M | 128K | $2.00 | $10.00 | 2026-06-30 | Source as of 2026-08-10 |
| Notes Mid tier balancing speed and intelligence. The listed price is introductory pricing through 2026-08-31; standard pricing of $3 input / $15 output takes effect on 2026-09-01. | ||||||||
| Claude Haiku 4.5 | Anthropic | Small | 200K | 64K | $1.00 | $5.00 | — | Source as of 2026-08-10 |
| Notes The fastest model in the lineup, commonly compared against rivals' budget tiers. Its 200K context and 64K max output are confirmed in the official docs. It uses the older tokenizer (pre-Claude 4.6 generation). Release date not verified here. | ||||||||
| Gemini 3 | Flagship | 1M | 64K | — | — | 2025-11-18 | Source as of 2026-08-10 | |
| Notes Flagship announced 2025-11 with 1M input / 64K output tokens. The Gemini API pricing page (checked 2026-08-10) lists only the 2.5/3.1/3.5/3.6 series and still shows no per-token price for Gemini 3. | ||||||||
| Gemini 3.5 Flash | Mid-tier | — | — | $1.50 | $9.00 | — | Source as of 2026-08-10 | |
| Notes Described in the model list as the most intelligent model for sustained frontier performance on agentic and coding tasks. It remains on the pricing page alongside the newer Gemini 3.6 Flash. Context window and release date are not listed. | ||||||||
| Gemini 3.6 Flash | Mid-tier | — | — | $1.50 | $7.50 | — | Source as of 2026-08-10 | |
| Notes Current Flash model, described in the model list as Google's latest model balancing speed with intelligence. Output pricing undercuts 3.5 Flash ($9 to $7.50). Context window and release date are not listed on the official pages. | ||||||||
| Gemini 3.1 Pro Preview | Flagship | — | — | $2.00 | $12.00 | — | Source as of 2026-08-10 | |
| Notes The Pro-tier model (preview) with published per-token pricing on the Gemini API pricing page. The listed price applies to prompts of 200k tokens or fewer; above that it becomes $4 input / $18 output. Context window and release date are not listed. | ||||||||
| Gemini 3.5 Flash-Lite | Small | — | — | $0.30 | $2.50 | — | Source as of 2026-08-10 | |
| Notes The current Flash-Lite tier listed on the Gemini API pricing page, in the price band commonly compared against rivals' budget tiers. Context window and release date are not listed on the pricing page. | ||||||||
| Muse Spark 1.1 | Meta | Small | 1M | — | $1.25 | $4.25 | 2026-07-09 | Source as of 2026-07-09 |
| Notes Meta's first paid API coding model; pricing is not in Meta's official blog and comes from TechCrunch reporting, so treat with caution. | ||||||||
| Qwen3.8-Max | Alibaba | Flagship | 1M | — | $2.00 | $6.00 | 2026-08-03 | Source as of 2026-08-10 |
| Notes A 2.4T-parameter MoE with 95B active, announced as the first Qwen-Max-class model to get open weights. Price and context window are as listed on QwenCloud. Max output is left null: QwenCloud shows "131.1K" while a configuration example in the official blog lists maxTokens 65536, so no single figure could be confirmed. | ||||||||
| Grok 4.5 | SpaceXAI | Flagship | 500K | — | $2.00 | $6.00 | 2026-07-08 | Source as of 2026-08-10 |
| Notes Top model co-trained with Cursor, positioned as Opus-class at lower cost and higher speed; 500K context confirmed in the official docs. Once a prompt reaches 200k tokens, the whole request is billed at $4 input / $12 output. | ||||||||
| Grok 4.3 | SpaceXAI | Flagship | 1M | — | $1.25 | $2.50 | — | Source as of 2026-08-10 |
| Notes Priced below Grok 4.5 with double the context at 1M tokens. Once a prompt reaches 200k tokens, the whole request is billed at $2.50 input / $5.00 output. Release date and max output are not listed in the official docs. | ||||||||
| Grok 4.5 Fast | SpaceXAI | Flagship | — | — | $4.00 | $18.00 | 2026-07-08 | Source as of 2026-07-08 |
| Notes A faster variant of Grok 4.5 offered on Cursor at a higher price; context window not published. | ||||||||
| Kimi K3 | Moonshot AI | Flagship | 1.0M | — | $3.00 | $15.00 | 2026-07-16 | Source as of 2026-08-02 |
| Notes A 2.8T-parameter MoE. Full weights were published on Hugging Face on 2026-07-27 as promised. Input price is $3 on cache miss and $0.30 on cache hit. The 1,048,576-token context is the figure stated on the model card (huggingface.co/moonshotai/Kimi-K3). | ||||||||
| Kimi K2 | Moonshot AI | Flagship | 128K | — | — | — | 2025-07-11 | Source as of 2025-07-11 |
| Notes Open-source 1T-parameter MoE; pricing is quoted in CNY (¥4 input / ¥16 output per 1M tokens) with no official USD rate. Figures date from 2025-07 and need refresh. | ||||||||
| Grok 4 | SpaceXAI | Flagship | 256K | — | — | — | 2025-07-09 | Source as of 2026-08-10 |
| Notes The previous-generation flagship before Grok 4.5 (released under the xAI name at the time), a multimodal model with a 256K context. As of 2026-08-10 it still does not appear in the official model list, and its per-token API price is not verified here. | ||||||||
No models match the current filters.