This Week in AI (July 25-31, 2026): Claude Opus 5, an 80% GPT-5.6 Price Cut, and a Compute-Stack Shake-Up

Claude Opus 5 lands at Opus 4.8 prices, OpenAI slashes GPT-5.6 Luna by 80%, SSI-NVIDIA and Nscale-Anyscale redraw the infrastructure map, and the open-weights debate heats up — this week's 11 stories in context.

This Week in AI (July 25-31, 2026): Claude Opus 5, an 80% GPT-5.6 Price Cut, and a Compute-Stack Shake-Up

The main character of this week (July 25-31) was price. Anthropic shipped Claude Opus 5 claiming top-tier performance at the same price as its predecessor, OpenAI cut GPT-5.6 — barely three weeks old — by up to 80%, and Cursor introduced its first regional pricing with an India-only plan at about $7 a month. With raw capability racing having settled into a rhythm, the competition is shifting to how cheaply the same capability can be delivered.

Beneath that, compute infrastructure kept consolidating. Sutskever’s SSI unveiled a long-term NVIDIA partnership aimed at scaling its compute by an order of magnitude, Recursive signed a $410 million compute contract with AWS, and Britain’s Nscale moved to buy Anyscale — the company behind Ray — reaching beyond GPUs and power into the software layer. The race to secure compute is becoming a race to stack it.

The tug-of-war over open weights ran in parallel: Moonshot released Kimi K3’s full weights on schedule, while Anthropic published an official position stating it has never advocated a ban, narrowing its regulatory asks to three points.

Model Updates: Stronger at the Same Price, Cheaper for the Same Model

On July 24, Anthropic announced Claude Opus 5. Pricing stays at $5 input / $25 output per million tokens, unchanged from Opus 4.8, while the company claims results matching — and on some benchmarks exceeding — its top model Claude Fable 5. It was available on all platforms from day one, making it a no-extra-cost upgrade for current Opus users.

On July 30, OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%. Luna now costs $0.20 input / $1.20 output per million tokens, and the API’s priority processing was replaced by Fast mode, which for Sol delivers up to 2.5× faster speeds at twice the price. For workloads that run small models around the clock, the monthly bill drops to one-fifth.

Compute Infrastructure: From Securing It to Stacking It

On July 27, Ilya Sutskever’s SSI announced a long-term strategic partnership with NVIDIA, gaining access to the next-generation Vera Rubin GPU platform to scale its compute by an order of magnitude; TechCrunch put the investment in the multiple billions and Bloomberg reported $5 billion. The next day, self-improving-AI lab Recursive signed a $410 million multi-year compute deal with AWS — most of the $650 million it raised in May committed to a single compute contract.

Then on July 30, British neocloud Nscale announced a definitive agreement to acquire Anyscale. Terms are officially undisclosed, though Bloomberg reported $1.65 billion. A GPU cloud absorbing the company founded by Ray’s creators signals that neoclouds’ pitch is shifting from “cheap GPUs” to “a full stack including operations software.”

Open Weights: Those Who Publish and Those Who Draw Lines

Moonshot AI released Kimi K3’s full weights on Hugging Face, meeting its self-imposed deadline. At 2.8 trillion total parameters, the company bills it as the world’s first open 3T-class model. Meanwhile on July 27, Anthropic published an official position on open-weights models under CEO Dario Amodei’s name, stating it has never advocated a ban and narrowing its asks to three: chip export controls on China, action against industrial-scale distillation, and safety testing for sufficiently capable models.

Specialized Models and Physical AI: The Next Battleground After General-Purpose

On July 27, Microsoft announced MAI-Cyber-1-Flash, its first security-specialized model, together with the agentic security system Project Perception — a notable break from repurposing general-purpose models for security work.

On July 30, Google DeepMind announced the Gemini Robotics 2 family, extending control from a humanoid’s upper body to whole-body motion including walking, and adding multi-robot collaboration. The reasoning model ER 2 is open to developers via the Gemini API and Google AI Studio.

AI Coding Tools: Localized Pricing and Buying the Conversation

On July 28, Cursor announced Cursor Start, an India-only plan at Rs649 (about $7) per month — its first regional pricing, landing ahead of the planned $60 billion SpaceX acquisition. And Devin maker Cognition announced its acquisition of The Interaction Company, the startup behind personal AI agent Poke, reportedly for low nine figures — a sign the agent race now extends to conversational experience itself.

Looking Ahead

Price cuts and infrastructure consolidation both reshape the cost structure of using AI. August brings the start of the EU AI Act’s transparency obligations, and the question of whether vertical integration like Nscale-Anyscale spreads to other neoclouds. We also track model price changes in our AI models observatory.

Wain publishes a digest like this once a week. Subscribe to the RSS feed to catch the weekly roundup along with our daily coverage.

We publish the latest AI news every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →