AI Model Pricing & Specs Tracker
A running comparison of pricing ($/1M tokens), context windows, and release dates for the API models offered by major AI labs — OpenAI, Anthropic, Google, xAI and more. Every row links to its source and the date it was last verified. Updated weekly.
Last updated: 2026-09-30
API Cost Calculator
Enter your expected monthly input and output tokens to see an estimated monthly cost for every model with published pricing, ranked from cheapest to most expensive. As a rule of thumb, 1 million tokens is roughly 750,000 English words.
| Model | Provider | Estimated monthly cost |
|---|---|---|
| GPT-6 Luna | OpenAI | $2.00 |
| GPT-5.6 Luna | OpenAI | $4.40 |
| DeepSeek-Flash | DeepSeek | $5.40 |
| DeepSeek-V4-Flash | DeepSeek | $7.04 |
| Gemini 3.5 Flash-Lite | $8.00 | |
| Hy4 preview | Tencent | $13.34 |
| Gemini 3.7 Flash | $15.00 | |
| Gemini 3.6 Flash | $15.00 | |
| Gemini 3.8 Flash | $15.00 | |
| Grok 4.3 | SpaceXAI | $17.50 |
| Claude Haiku 4.5 | Anthropic | $20.00 |
| Muse Spark 1.1 | Meta | $21.00 |
| Muse Spark 1.2 | Meta | $21.00 |
| Muse Spark 1.3 | Meta | $21.00 |
| DeepSeek-V4-Pro | DeepSeek | $21.12 |
| Qwen3.8-Max | Alibaba | $32.00 |
| Grok 4.6 | SpaceXAI | $32.00 |
| Grok 4.5 | SpaceXAI | $32.00 |
| Grok 4.7 | SpaceXAI | $32.00 |
| Gemini 3.5 Flash | $33.00 | |
| Claude Sonnet 5.5 | Anthropic | $40.00 |
| Claude Sonnet 5 | Anthropic | $40.00 |
| GPT-6 Sol | OpenAI | $40.00 |
| GPT-6.1 Sol | OpenAI | $40.00 |
| GPT-5.6 Terra | OpenAI | $44.00 |
| Gemini 3.1 Pro Preview | $44.00 | |
| Kimi K3 | Moonshot AI | $60.00 |
| Grok 4.5 Fast | SpaceXAI | $76.00 |
| GPT-5.6 Sol | OpenAI | $80.00 |
| Claude Opus 5.5 | Anthropic | $80.00 |
| Claude Opus 5 | Anthropic | $100.00 |
| Claude Opus 4.8 | Anthropic | $100.00 |
| Claude Opus 4.7 | Anthropic | $100.00 |
| GPT-6 Astra | OpenAI | $200.00 |
| Claude Fable 5 | Anthropic | $200.00 |
| Claude Mythos 5 | Anthropic | $200.00 |
| Claude Fable 5.1 | Anthropic | $200.00 |
| Claude Mythos 5.1 | Anthropic | $200.00 |
Excluded due to unpublished pricing: 3 model(s)
* Rough estimate only — does not account for prompt caching or batch discounts. See each provider's source link for exact pricing.
All models
| Model | Provider | Tier | Context | Max output | Input ($/1M tokens) | Output ($/1M tokens) | Released | Source |
|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | Flagship | — | — | $4.00 | $20.00 | 2026-07-09 | Source as of 2026-09-28 |
| Notes Flagship of the GPT-5.6 family. On 2026-08-21 pricing dropped from $5 to $4 input (-20%) and $30 to $20 output (-33%); the official changelog calls this promotional pricing available at least through 2026-11-21. Long context (over 272K tokens) is $8 input / $30 output, and Batch is $2 input / $10 output. Fast mode is billed at double the standard rate. Context window and max output are not listed on the official pricing page. | ||||||||
| GPT-5.6 Terra | OpenAI | Mid-tier | — | — | $2.00 | $12.00 | 2026-07-09 | Source as of 2026-09-28 |
| Notes Balanced everyday-work tier; price cut 20% on 2026-07-30. Fast mode is billed at double the standard rate. Context window and max output are not listed on the official pricing page. As of 2026-09-28 this model no longer appears in the text-model table on the official pricing page (developers.openai.com/api/docs/pricing); the table lists gpt-6-astra, gpt-6-sol, gpt-6-luna, gpt-5.6-sol and gpt-5.6-cyber. The prices kept here are the last confirmed listing, and current availability is unverified in this dataset. | ||||||||
| GPT-5.6 Luna | OpenAI | Small | — | — | $0.20 | $1.20 | 2026-07-09 | Source as of 2026-09-28 |
| Notes Most cost-efficient tier; price cut 80% on 2026-07-30. Fast mode is billed at double the standard rate. Context window and max output are not listed on the official pricing page. As of 2026-09-28 this model no longer appears in the text-model table on the official pricing page. Its successor gpt-6-luna is listed at $0.10 input / $0.50 output (separate row). The prices kept here are the last confirmed listing. | ||||||||
| GPT-6 Astra | OpenAI | Flagship | — | — | $10.00 | $50.00 | 2026-09-03 | Source as of 2026-09-28 |
| Notes Flagship of the GPT-6 generation, released on 2026-09-03 and offered through the OpenAI API and Amazon Bedrock. Standard pricing is $10 input / $50 output, with cached input at $1.00. Fast mode is billed at double the standard rate ($20 input / $100 output). That is 2.5x the input and output rates of GPT-5.6 Sol. Context window and max output are not listed on the official pricing page. | ||||||||
| Claude Fable 5 | Anthropic | Flagship | 1M | 128K | $10.00 | $50.00 | 2026-06-09 | Source as of 2026-09-28 |
| Notes Anthropic's most capable widely released model, GA on 2026-06-09. It uses the tokenizer introduced with Opus 4.7, which produces roughly 30% more tokens for the same text — worth noting when comparing costs. Price and specs confirmed in the official docs. It remains listed alongside the successor Fable 5.1 released on 2026-09-01; base rates are identical but cache reads cost $1 versus $0.25 on 5.1. | ||||||||
| Claude Mythos 5 | Anthropic | Flagship | 1M | 128K | $10.00 | $50.00 | 2026-06-09 | Source as of 2026-09-28 |
| Notes Shares Fable 5's specs and pricing; offered under limited availability for defensive cybersecurity work through Project Glasswing, by invitation only with no self-serve sign-up. It remains listed alongside the successor Mythos 5.1 released on 2026-09-01; cache reads cost $1 versus $0.25 on 5.1. | ||||||||
| Claude Opus 5 | Anthropic | Flagship | 1M | 128K | $5.00 | $25.00 | 2026-07-24 | Source as of 2026-09-28 |
| Notes New-generation Opus at the same price as Opus 4.8. Fast mode (research preview) is $10 input / $50 output, double the standard rate. Note that the Opus 4.7-era tokenizer produces roughly 30% more tokens for the same text. | ||||||||
| Claude Opus 4.8 | Anthropic | Flagship | 1M | 128K | $5.00 | $25.00 | — | Source as of 2026-09-28 |
| Notes A key baseline model across vendor benchmarks, now listed under legacy models in the official docs. Fast mode (research preview) is $10 input / $50 output. Release date not verified here. | ||||||||
| Claude Opus 4.7 | Anthropic | Flagship | 1M | 128K | $5.00 | $25.00 | — | Source as of 2026-09-28 |
| Notes Prior-generation flagship, listed under legacy models in the official docs. It introduced the newer tokenizer that produces roughly 30% more tokens for the same text. Fast mode is not supported (requests with speed: "fast" return an error). Release date not verified here. | ||||||||
| Claude Sonnet 5.5 | Anthropic | Mid-tier | 1M | 128K | $2.00 | $10.00 | 2026-09-28 | Source as of 2026-09-29 |
| Notes The second model in the Claude 5.5 family. The official pricing page lists $2 input / $10 output, cache writes of $2.50 (5m) and $4 (1h), and $0.20 cache reads. Unit prices match Sonnet 5; Anthropic says cost per task is up to 30% lower in its own testing because fewer tokens are needed for the same work, rather than because rates were cut. The official model list gives a 1M-token context, 128K max output, a default effort of medium and the model ID claude-sonnet-5-5. As of 2026-09-28 the main pricing table lists this model in place of Sonnet 5, which moved to the Additional models section at unchanged prices. | ||||||||
| Claude Sonnet 5 | Anthropic | Mid-tier | 1M | 128K | $2.00 | $10.00 | 2026-06-30 | Source as of 2026-09-28 |
| Notes Mid tier balancing speed and intelligence. The $2 input / $10 output rate, originally introductory pricing through 2026-08-31, is now the standard price; the official pricing page states the increase to $3/$15 scheduled for 2026-09-01 will not occur. | ||||||||
| Claude Haiku 4.5 | Anthropic | Small | 200K | 64K | $1.00 | $5.00 | — | Source as of 2026-09-28 |
| Notes The fastest model in the lineup, commonly compared against rivals' budget tiers. Its 200K context and 64K max output are confirmed in the official docs. It uses the older tokenizer (pre-Claude 4.6 generation). Release date not verified here. | ||||||||
| Claude Fable 5.1 | Anthropic | Flagship | 1M | — | $10.00 | $50.00 | 2026-09-01 | Source as of 2026-09-28 |
| Notes Successor to Fable 5 and the top generally available model. Base input and output rates match Fable 5, but cache reads are priced at 0.025x the base input rate ($0.25/MTok) instead of the 0.1x multiplier used by other models. Five-minute cache writes are $12.50, one-hour cache writes $20, and Batch is $5 input / $25 output. The 1M context figure comes from the Claude Code release notes; max output is not listed on the official pricing page. | ||||||||
| Claude Mythos 5.1 | Anthropic | Flagship | — | — | $10.00 | $50.00 | 2026-09-01 | Source as of 2026-09-28 |
| Notes The same underlying model as Fable 5.1 with different safeguard levels, offered under limited availability as noted on the official pricing table. Rates and cache pricing match Fable 5.1, including $0.25/MTok cache reads. Access is limited to vetted cyberdefenders and life scientists, initially at a set of US organizations. Context window and max output are not listed on the official pricing page. | ||||||||
| Gemini 3 | Flagship | 1M | 64K | — | — | 2025-11-18 | Source as of 2026-09-14 | |
| Notes Flagship announced 2025-11 with 1M input / 64K output tokens. The Gemini API pricing page (checked 2026-08-10) lists only the 2.5/3.1/3.5/3.6 series and still shows no per-token price for Gemini 3. | ||||||||
| Gemini 3.5 Flash | Mid-tier | — | — | $1.50 | $9.00 | — | Source as of 2026-09-28 | |
| Notes Described in the model list as the most intelligent model for sustained frontier performance on agentic and coding tasks. It remains on the pricing page alongside the newer Gemini 3.6 Flash. Context window and release date are not listed. | ||||||||
| Gemini 3.7 Flash | Mid-tier | — | — | $0.75 | $3.75 | 2026-08-13 | Source as of 2026-09-28 | |
| Notes Main model for coding and agents. The official pricing page states $0.75/$3.75 applies through 2026-12-31, with $1.50/$7.50 from 2027-01-01. Listing on the official Gemini API pricing page confirmed as of 2026-08-24. Context window and max output are not listed. | ||||||||
| Gemini 3.6 Flash | Mid-tier | — | — | $0.75 | $3.75 | — | Source as of 2026-09-28 | |
| Notes Repriced to the same introductory rate as Gemini 3.7 Flash. The official pricing page states $0.75/$3.75 applies through 2026-12-31 and reverts to $1.50/$7.50 on 2027-01-01 (as of 2026-08-17 it was $1.50/$7.50). Context window and release date are not listed on the official pages. | ||||||||
| Gemini 3.1 Pro Preview | Flagship | — | — | $2.00 | $12.00 | — | Source as of 2026-09-28 | |
| Notes The Pro-tier model (preview) with published per-token pricing on the Gemini API pricing page. The listed price applies to prompts of 200k tokens or fewer; above that it becomes $4 input / $18 output. Context window and release date are not listed. | ||||||||
| Gemini 3.5 Flash-Lite | Small | — | — | $0.30 | $2.50 | — | Source as of 2026-09-28 | |
| Notes The current Flash-Lite tier listed on the Gemini API pricing page, in the price band commonly compared against rivals' budget tiers. Context window and release date are not listed on the pricing page. | ||||||||
| Gemini 3.8 Flash | Mid-tier | — | — | $0.75 | $3.75 | 2026-09-02 | Source as of 2026-09-28 | |
| Notes The newest Flash model, launched at the same introductory price as 3.7 Flash. A footnote on the official pricing page states that $0.75/$3.75 applies through 2026-12-31 and that $1.50/$7.50 takes effect on 2027-01-01. Pricing for the vetted-access cyber variant, 3.8 Flash Cyber, has not been published. Context window and max output are not listed on the pricing page. | ||||||||
| Muse Spark 1.1 | Meta | Small | 1M | — | $1.25 | $4.25 | 2026-07-09 | Source as of 2026-09-14 |
| Notes Meta's first paid API coding model; pricing confirmed on the official Meta Model API pricing page (Standard tier, $0.15 cached input). Previously sourced from TechCrunch reporting, now replaced with the official figures. | ||||||||
| Muse Spark 1.2 | Meta | Small | — | — | $1.25 | $4.25 | 2026-08-05 | Source as of 2026-09-14 |
| Notes Coding-focused model released 2026-08-05. Standard tier pricing matches 1.1 ($0.15 cached input). The Contributor tier (muse-spark-1.2-contributor), which permits prompts and completions to train future Meta models, is far cheaper at $0.10 input / $0.20 output. Context window and max output are not listed on the official pricing page. | ||||||||
| Muse Spark 1.3 | Meta | Small | — | — | $1.25 | $4.25 | — | Source as of 2026-09-14 |
| Notes The latest Muse Spark release. Standard-tier rates are unchanged from 1.1 and 1.2. A cheaper Contributor tier ($0.10 input / $0.20 output) is offered in exchange for permission to use prompts and completions to train future Meta models. Context window and release date are not listed on the official pricing page. | ||||||||
| Qwen3.8-Max | Alibaba | Flagship | 1M | — | $2.00 | $6.00 | 2026-08-03 | Source as of 2026-09-07 |
| Notes A 2.4T-parameter MoE with 95B active, announced as the first Qwen-Max-class model to get open weights. Price and context window are as listed on QwenCloud. Max output is left null: QwenCloud shows "131.1K" while a configuration example in the official blog lists maxTokens 65536, so no single figure could be confirmed. | ||||||||
| Grok 4.6 | SpaceXAI | Flagship | 500K | — | $2.00 | $6.00 | — | Source as of 2026-09-28 |
| Notes Current top model, listed first in the official docs model list; same price and same 500K context as Grok 4.5. Once a prompt reaches 200k tokens, the whole request is billed at $4 input / $12 output. The official docs give a knowledge cut-off of 2026-02-01. Release date and max output are not listed. As of 2026-09-28 the only model listed in the official docs model list (docs.x.ai/developers/models) is Grok 4.7; this model no longer appears. Grok 4.7 is served at the same price ($2 input / $6 output) and the same 500K context (separate row). The prices kept here are the last confirmed listing. | ||||||||
| Grok 4.5 | SpaceXAI | Flagship | 500K | — | $2.00 | $6.00 | 2026-07-08 | Source as of 2026-09-28 |
| Notes Top model co-trained with Cursor, positioned as Opus-class at lower cost and higher speed; 500K context confirmed in the official docs. Once a prompt reaches 200k tokens, the whole request is billed at $4 input / $12 output. Still absent from the official docs model list as of 2026-09-28 (the list shows only Grok 4.7), so current availability is unverified in this dataset. | ||||||||
| Grok 4.3 | SpaceXAI | Flagship | 1M | — | $1.25 | $2.50 | — | Source as of 2026-09-28 |
| Notes Priced below Grok 4.5 with double the context at 1M tokens. Once a prompt reaches 200k tokens, the whole request is billed at $2.50 input / $5.00 output. Release date and max output are not listed in the official docs. Still absent from the official docs model list as of 2026-09-28 (the list shows only Grok 4.7), so current availability is unverified in this dataset. | ||||||||
| Grok 4.5 Fast | SpaceXAI | Flagship | — | — | $4.00 | $18.00 | 2026-07-08 | Source as of 2026-08-24 |
| Notes Faster variant of Grok 4.5 offered through Cursor, priced above the base model; context window unconfirmed officially. As of 2026-08-24 no matching model ID appears in the official docs model list, so current availability is unverified in this dataset. | ||||||||
| Kimi K3 | Moonshot AI | Flagship | 1.0M | — | $3.00 | $15.00 | 2026-07-16 | Source as of 2026-09-14 |
| Notes A 2.8T-parameter MoE. Full weights were published on Hugging Face on 2026-07-27 as promised. Input price is $3 on cache miss and $0.30 on cache hit. The 1,048,576-token context is the figure stated on the model card (huggingface.co/moonshotai/Kimi-K3). | ||||||||
| Kimi K2 | Moonshot AI | Flagship | 128K | — | — | — | 2025-07-11 | Source as of 2025-07-11 |
| Notes Open-source 1T-parameter MoE; pricing is quoted in CNY (¥4 input / ¥16 output per 1M tokens) with no official USD rate. Figures date from 2025-07 and need refresh. | ||||||||
| Grok 4 | SpaceXAI | Flagship | 256K | — | — | — | 2025-07-09 | Source as of 2026-08-24 |
| Notes Previous-generation flagship before Grok 4.5, released under the xAI name; a multimodal model with a 256K context. Still absent from the official model list as of 2026-08-24, so API usage pricing remains unverified in this dataset. | ||||||||
| DeepSeek-V4-Pro | DeepSeek | Flagship | 1M | 384K | $1.32 | $3.96 | 2026-08-13 | Source as of 2026-09-28 |
| Notes Switched to time-of-day pricing at 16:00 UTC on 2026-08-16. The values here are peak-hour (01:00-04:00 and 06:00-10:00 UTC) cache-miss input / output rates; off-peak rates are half. Off-peak is $0.66 input / $1.98 output. Cache-hit input is $0.044 at peak and $0.022 off-peak. The pricing page lists the version as DeepSeek-V4-Pro-0813. As of 2026-09-14 the pricing page carries a note saying that, in response to user demand, API service for DeepSeek V4 Pro continues beyond September 14, 2026 with the billing method unchanged. | ||||||||
| DeepSeek-V4-Flash | DeepSeek | Mid-tier | 1M | 384K | $0.44 | $1.32 | 2026-07-31 | Source as of 2026-09-14 |
| Notes Switched to time-of-day pricing at 16:00 UTC on 2026-08-16. The values here are peak-hour (01:00-04:00 and 06:00-10:00 UTC) cache-miss input / output rates; off-peak rates are half. Off-peak is $0.22 input / $0.66 output. Cache-hit input is $0.014 at peak and $0.007 off-peak. The pricing page lists the version as DeepSeek-V4-Flash-0731. As of 2026-09-14 this model no longer appears on the pricing page: the names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but the page notes that the corresponding models have been retired and that such requests are served by DeepSeek-V4.1-Flash and billed at the Flash price. The prices kept here are the pre-retirement figures; see the "DeepSeek-Flash" row for current rates. | ||||||||
| DeepSeek-Flash | DeepSeek | Mid-tier | 1M | 384K | $0.30 | $1.20 | — | Source as of 2026-09-28 |
| Notes The current Flash-line model on the pricing page, succeeding DeepSeek-V4-Flash. The API model name is deepseek-flash and the page lists the version as DeepSeek-V4.1-Flash. The values here are peak-hour (01:00-04:00 and 06:00-10:00 UTC, Monday to Friday) cache-miss input / output rates; off-peak is half ($0.15 input / $0.6 output). Cache-hit input is $0.006 at peak and $0.003 off-peak. Unlike V4-Pro it supports image input. The release date is not stated on the official page. | ||||||||
| Hy4 preview | Tencent | Flagship | 1M | — | $0.83 | $2.50 | 2026-08-28 | Source as of 2026-08-31 |
| Notes An open-weights MoE with 770B total parameters and 49B activated per token, published under Apache 2.0. Cache-hit input is $0.042. The announcement lists a single "API pricing" figure without breaking it out by provider (Tencent Cloud TokenHub / OpenRouter). The 1M context length comes from the model card; max output is not stated officially. | ||||||||
| Claude Opus 5.5 | Anthropic | Flagship | 1M | 128K | $4.00 | $20.00 | 2026-09-22 | Source as of 2026-09-29 |
| Notes The current top model in the Opus line. The official pricing page lists $4 input / $20 output, $5 cache writes and $0.20 cache reads. Against Opus 5 ($5 input / $25 output, $0.50 cache reads) that is 20% lower on input and output and 60% lower on cache reads. A fast mode with up to 2.5x speed is available in Claude Code and the Claude Platform at $8 input / $40 output. The official model list gives a 1M-token context. The official model list gives a 128K max output (confirmed 2026-09-29). | ||||||||
| GPT-6 Sol | OpenAI | Mid-tier | — | — | $2.00 | $10.00 | 2026-09-22 | Source as of 2026-09-28 |
| Notes The mid tier of the GPT-6 generation. The official pricing page lists $2 input / $10 output with $0.20 cached input. That is 50% below GPT-5.6 Sol ($4 input / $20 output), though that $4/$20 is itself promotional pricing running through 2026-11-21. An OpenAI spokesperson was reported by VentureBeat to have said the new Sol and Luna rates are permanent rather than promotional or introductory. Fast mode is billed at double the standard rate ($4 input / $20 output). Context window and max output are not listed on the pricing page. | ||||||||
| GPT-6.1 Sol | OpenAI | Mid-tier | 1.1M | 128K | $2.00 | $10.00 | 2026-09-29 | Source as of 2026-09-30 |
| Notes Released at DevDay 2026 on 2026-09-29 as an upgrade to GPT-6 Sol. OpenAI positions it as nearly matching GPT-6 Astra's intelligence on agentic coding, computer use and professional work at one-fifth of Astra's standard input and output token prices. Input at $2 and output at $10 match GPT-6 Sol; only cached input changed, from $0.20 to $0.10 (5% of the uncached input rate). Cache writes are $2.50. Prompts above 272K tokens are billed at 2x for input and cache rates and 1.5x for output. reasoning.effort supports low/medium/high/xhigh/max; none and minimal, which GPT-6 Sol accepted, are not supported. Tool calling requires the Responses API (Chat Completions is supported without tools). Knowledge cutoff is 2026-04-30 and maximum input is 922,000 tokens. US and EU data residency are supported, but Fast mode is unavailable with EU data residency. | ||||||||
| GPT-6 Luna | OpenAI | Small | — | — | $0.10 | $0.50 | 2026-09-22 | Source as of 2026-09-28 |
| Notes The smallest tier of the GPT-6 generation and the cheapest model in the current lineup. The official pricing page lists $0.10 input / $0.50 output with $0.01 cached input. It comes down from GPT-5.6 Luna ($0.20 input / $1.20 output); the official table says "50% cheaper", though the output reduction works out to roughly 58%. Fast mode is billed at double the standard rate. Context window and max output are not listed on the pricing page. | ||||||||
| Grok 4.7 | SpaceXAI | Flagship | 500K | — | $2.00 | $6.00 | 2026-09-21 | Source as of 2026-09-28 |
| Notes As of 2026-09-28 the only Grok model in the official docs model list. The docs confirm $2 input / $6 output and a 500K context. The announcement states that price and speed match Grok 4.6 and that a fast variant with double the output speed costs twice as much. The docs give a knowledge cut-off of May 2026. Release date and max output are not in the docs; the date here is taken from the publication date of the official announcement. | ||||||||
No models match the current filters.
How you may use this data
The pricing and specification data on this page is available under the Creative Commons Attribution 4.0 International license (CC BY 4.0) and may be reused, commercially or otherwise. Please credit the source with the site name, the URL of this page, and the date the data was confirmed.
Example citation: Wain, "AI Model Pricing & Specs Comparison," confirmed 2026-09-30, https://ai.wain.blog/en/ai-models/