Google Introduces Gemini 2.5 as a 'Thinking Model' — With Parallel-Thinking Deep Think Added Later

Google announced Gemini 2.5 as a 'thinking model' that reasons before responding in March 2025 (SWE-Bench Verified 63.8%), then later shipped the parallel-thinking 'Deep Think' feature separately. What the official announcements actually said.

Google Introduces Gemini 2.5 as a 'Thinking Model' — With Parallel-Thinking Deep Think Added Later

On March 25, 2025, Google announced Gemini 2.5 as a “thinking model”1. Rather than returning an answer instantly, it reasons through its own thoughts before responding, which Google says improves performance and accuracy. The background of this “think before answering” approach is covered in what a reasoning model is.

Note (updated July 2026): This article originally contained errors about the timing and performance figures. It has been fully revised against the official announcements. In particular, “Deep Think” was not introduced together with Gemini 2.5 in March 2025, but shipped later as a separate feature.

Gemini 2.5: a “thinking model” that reasons before responding

Google said it combined a significantly enhanced base model with improved post-training, and that it is building these thinking capabilities directly into all of its models1. Its previous thinking model was “Gemini 2.0 Flash Thinking.” Because a thinking model reasons internally before generating a response, it tends to be more accurate on tasks that require multi-step logic or mathematical and scientific reasoning.

Benchmarks and availability at launch

At launch, Gemini 2.5 Pro was said to top the LMArena leaderboard by a significant margin, and to lead on math and science benchmarks such as GPQA and AIME 20251. On the difficult Humanity’s Last Exam benchmark it scored 18.8%, and on the software-engineering benchmark SWE-Bench Verified it scored 63.8%1.

On availability, it launched in Google AI Studio and in the Gemini app (for Gemini Advanced users), with Vertex AI support to follow. The context window started at 1 million tokens, with 2 million tokens announced to come1.

Deep Think, the parallel-thinking feature added later

“Deep Think” was not a feature introduced alongside the Gemini 2.5 announcement. Google unveiled Deep Think at Google I/O 2025 and made it generally available to Google AI Ultra subscribers on August 1, 20252.

What distinguishes Deep Think is that, rather than laying out a thought process sequentially, it uses “extended, parallel thinking” together with novel reinforcement-learning techniques2. In this approach, Gemini generates many ideas at once and considers them simultaneously, revising or combining different ideas over time. It is less about following a single line of thought and more about running multiple hypotheses in parallel to compare and integrate them.

Deep Think’s performance and access

According to Google, Deep Think reaches state-of-the-art performance on the coding benchmark LiveCodeBench V6 and on Humanity’s Last Exam, and the standard version reached Bronze level on the 2025 International Mathematical Olympiad (IMO) benchmark2. A specialized version was reported to achieve gold-medal standard at the 2025 IMO. Note, however, that the gold-medal result is the achievement of a specialized version, under different conditions from the generally available standard Deep Think.

Access was limited to Google AI Ultra subscribers, who toggle it within the Gemini app2. Its intended uses include creative problem-solving, coding, scientific discovery, and iterative design that require advanced reasoning.

Weighing Thinking Depth Against Speed and Cost — and What Came Next

Gemini 2.5 and Deep Think illustrate how AI is evolving in the direction of controlling “how much a model thinks before responding.” Thinking models tend to improve accuracy on reasoning-heavy tasks, but they also tend to increase response time and compute cost. That Deep Think was offered only on the higher AI Ultra subscription reflects the cost structure of such features.

Weighing “depth of thinking” against “speed and cost” is a practical consideration when choosing a model. For routine processing, a lightweight model is often sufficient; using thinking models or parallel-thinking features only where complex reasoning is genuinely needed is a realistic way to divide usage.

The Gemini series has continued through further generations, and the successor Gemini 3 strengthens reasoning performance further. Gemini 2.5 and Deep Think covered here can be seen as the starting point of the “thinking model” design philosophy that underpins them.

Sources

  1. Gemini 2.5: Our most intelligent AI model - Google Official Blog (March 25, 2025)
  2. Try Deep Think in the Gemini app - Google Official Blog (August 1, 2025, Deep Think general availability)

We publish the latest AI news every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →