OpenAI o3 Model Achieves Historic Breakthrough on AGI Benchmark - 87.5% on ARC-AGI Shocks AI Industry

OpenAI's o3 model announced in December 2024 achieved 87.5% on the ARC-AGI benchmark, a key indicator of AGI progress. Exploring revolutionary deliberative alignment safety features and computational cost challenges.

OpenAI o3 Model Achieves Historic Breakthrough on AGI Benchmark - 87.5% on ARC-AGI Shocks AI Industry

On December 20, 2024, OpenAI announced the next-generation reasoning model “o3” on the final day of their “12 Days of Shipmas” event. This model achieved a historic score of 87.5% on the ARC-AGI benchmark, a key indicator of AGI (Artificial General Intelligence) progress, sending shockwaves through the AI industry.

Strategic Naming: Skipping o2

Interestingly, OpenAI named the successor to o1 as “o3” rather than “o2.” This was a strategic decision to avoid trademark issues with UK telecommunications company O2. The o3 was announced alongside o3-mini, a lightweight version, forming two model families optimized for specific tasks.

Historic Breakthrough on ARC-AGI Benchmark

The most notable achievement of o3 is its remarkable performance on the ARC-AGI benchmark. This benchmark is a challenging test designed to measure AI’s true reasoning capabilities, where previous models showed minimal progress.

GPT-3 in 2020 scored 0%, GPT-4o in 2024 only managed 5%. Claude 3.5 Sonnet achieved 14%, and o1-preview reached 18%. However, o3 achieved a dramatic improvement with 75.7% in standard compute settings ($10k compute limit) and 87.5% in low compute settings.

The speed of this progress has surprised AI researchers. The improvement from 0% to 87.5% in just four years demonstrates the exponential evolution of AI technology. However, it was revealed that running the high compute setting (172 times the standard computation) would cost an estimated $1.14 million (approximately 170 million yen) to solve all 400 public test problems.

Revolutionary “Deliberative Alignment” Technology

The core technology supporting o3’s safety is “deliberative alignment.” This is a groundbreaking approach where the model “thinks” by referencing OpenAI’s safety policies before generating responses.

Traditional AI models relied on safety measures embedded during pre-training. However, o3 dynamically judges safety during inference. Specifically, when faced with user questions, it first internally references safety policy text, then analyzes the hidden intent and potential dangers of the question through a “chain of thought” process.

For example, when asked about creating fraudulent documents, o3 internally quotes OpenAI policies, identifies that the request aims for illegal activities, and appropriately refuses to respond. This technology makes o3 one of OpenAI’s safest models.

Outstanding Benchmark Performance

o3 has achieved excellent results across many benchmarks beyond ARC-AGI.

In mathematics, it achieved a 96.7% score on the 2024 American Invitational Mathematics Examination (AIME), with only one incorrect answer—a remarkable result. In science, it achieved 87.7% on GPQA Diamond, a graduate-level biology, physics, and chemistry problem set.

Particularly noteworthy is the 25.2% score on EpochAI’s Frontier Math benchmark. While other models couldn’t exceed 2%, o3 demonstrated more than 10 times better performance. Coding capabilities also improved significantly, achieving a Codeforces rating of 2727 and outperforming o1 by 22.8 points on SWE-Bench Verified.

Challenges and Expectations for Practical Implementation

While o3’s performance is revolutionary, practical implementation faces several challenges. The biggest challenge is computational cost. High compute settings (172 times the computation) required to achieve high performance would cost an estimated $1.14 million to solve all 400 public test problems. Even efficient settings cost about $6,677 (approximately 1 million yen), which is expensive for general use.

Additionally, the ARC Prize organization has expressed caution, stating that “passing the ARC-AGI test does not mean achieving AGI.” They point out that o3 still fails at very simple tasks and fundamental differences exist compared to human intelligence.

Future Developments

OpenAI accepted applications for early access from safety and security researchers until January 10, 2025. General release is planned for early 2025, but specific dates and pricing structures have not yet been announced.

The AI industry has entered a new competitive phase with o3’s emergence. How companies respond to this technological breakthrough and when practical AGI might be realized are developments worth watching.

Sources

We publish the latest AI news every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →