OpenAI Cuts GPT-5.6 Luna Prices by 80% and Terra by 20%, Replaces Priority Processing With Fast Mode

OpenAI cut GPT-5.6 Luna to $0.20/$1.20 per million tokens (80% off) and Terra by 20% on July 30, and replaced the API's Priority Processing with Fast mode — just three weeks after launch.

OpenAI Cuts GPT-5.6 Luna Prices by 80% and Terra by 20%, Replaces Priority Processing With Fast Mode

Starting July 30, 2026, OpenAI cut the API price of GPT-5.6 Luna, the smallest model in the family, by 80%, and mid-tier Terra by 20%1. The revision comes just three weeks after the general availability launch on July 9, while pricing for the flagship Sol stays unchanged23.

Alongside the cuts, the API’s Priority Processing option has been replaced by “Fast mode.” For Sol, Fast mode delivers up to 2.5× faster speeds than standard processing at twice the price, and the change is backward compatible: requests tagged priority automatically move to Fast mode1.

The New Prices: Luna Drops to $0.20 per Million Input Tokens

The revised prices (per million tokens, short context) are as follows2. Launch prices are as of the July 9 release.

ModelAt launch (Jul 9)After the cut (Jul 30)Reduction
GPT-5.6 Sol$5.00 in / $30.00 outunchanged-
GPT-5.6 Terra$2.50 in / $15.00 out$2.00 in / $12.00 out20%
GPT-5.6 Luna$1.00 in / $6.00 out$0.20 in / $1.20 out80%

Long context carries separate, higher rates (for Luna, $0.40 input / $1.80 output)2. The pricing page lists Fast mode rates for all three tiers — Sol, Terra, and Luna — with Sol at $10.00 input / $60.00 output (the changelog specifies the 2.5× speedup for Sol)12.

OpenAI is reported to claim that a task that cost a dollar on models from a year ago now runs at about 6 cents on Luna, nearly nine times faster3. For workloads that keep a small model running around the clock — bulk document processing, classification, extraction — the monthly API bill literally drops to one-fifth.

The Money Comes From “the Model Improving Its Own Infrastructure”

According to THE DECODER, OpenAI attributes the cuts to infrastructure efficiency gains made using GPT-5.6 Sol itself: the company says it reduced GPU deployment costs by 20% and improved token generation by more than 15% through speculative decoding3. The claim is that the model optimizes the company’s own inference stack and the savings flow back into prices — if accurate, a sign that “AI lowering the cost of running AI” is becoming a weapon in the price war.

External pressure is also cited. The same article points to growing price pressure from low-cost Chinese providers and notes that Microsoft is promoting its own MAI models as alternatives to OpenAI3. Google, too, has just strengthened its token-efficient, low-price line with the Gemini 3.6 Flash series. Unit prices for small and mid-size models have become one of the most fiercely contested fronts among major providers this summer.

Living With API Price Sheets That Expire in Weeks

This revision is the kind where costs fall without switching models — you win by waiting. The flip side is that any pricing assumption can collapse within weeks. When comparing models for a use case or weighing tool subscriptions, record the date you checked alongside the numbers, or next month you may reach a different conclusion. Our own AI coding tools pricing comparison treats every price as of its verification date for the same reason.

How much you gain in practice depends on whether you have carved out the workloads a Luna-class model can handle — classification, summarization, routing. If you only use Sol, nothing changes. Conversely, for conversational products where latency is the bottleneck, paying double for Fast mode is now an option. Shipping a price cut and a speed tier at the same time reads as OpenAI formalizing the split between “cheap at scale” and “fast at a premium” as standard API features.

Sources

  1. Changelog - OpenAI official API changelog (entry dated July 30, 2026)
  2. Pricing - OpenAI official API pricing page
  3. OpenAI goes full China pricing mode with an 80 percent cut to its most affordable GPT-5.6 model - THE DECODER

We publish the latest AI news every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →