OpenAI officially released its most advanced reasoning AI model, “o3-pro,” on June 10, 20251. o3-pro serves as the top-tier model replacing the previous o1-pro model and is now available to ChatGPT Pro and Team users. Enterprise and Edu version users will gain access the following week.
o3-pro is positioned as the flagship version of the o3 family announced in December 2024. Leveraging the characteristics of a reasoning-based model that solves complex problems step-by-step, it has been confirmed to deliver high performance in specialized fields such as physics, mathematics, and programming. According to OpenAI’s internal testing, expert evaluations show that o3-pro outperforms o3 across all test categories, particularly excelling in science, education, programming, business, and writing assistance1.
Benchmark Performance Surpasses Google and Anthropic
Particularly noteworthy in o3-pro’s performance evaluation is its achievement of results that surpass major competing models. In the AIME 2024 benchmark evaluating mathematical ability, it recorded a score higher than Google’s top-tier model Gemini 2.5 Pro1. Furthermore, in the GPQA Diamond test evaluating PhD-level scientific knowledge, it demonstrated performance that even surpasses Anthropic’s recently released Claude 4 Opus1.
The overall performance improvement of the o3 family is also remarkable, with o3-mini’s high-precision mode achieving 96.7% accuracy, a significant improvement from o1’s 83.3%2. Particularly in mathematics, o3 recorded a score of 87.7% and demonstrated powerful reasoning capabilities in graduate-level biology, physics, and chemistry problems. Additionally, in the Frontier Math benchmark, while other models remained below 2%, o3 achieved a breakthrough result of 25.2%, according to OpenAI’s announcement2.
Diverse Tool Integration and Enhanced Practicality
A major characteristic of o3-pro lies in its multifunctionality that extends beyond mere language processing. The model provides access to a wide range of tools including web search, file analysis, visual input reasoning, Python execution, and personalized responses utilizing memory1. This enables the processing of more complex and practical tasks.
API pricing is set at $20 per million input tokens and $80 per million output tokens1, representing a pricing structure designed for enterprise-level utilization. Additionally, o3-mini achieves a 63% cost reduction compared to o1-mini, delivering 24% faster responses with 39% fewer critical errors, demonstrating improved cost performance2.
Current Limitations and Future Development
While o3-pro boasts high performance, several limitations exist. The most notable is response speed, with OpenAI reporting that processing takes longer compared to o1-pro1. Additionally, due to technical issues, temporary chat functionality in ChatGPT is currently disabled, and image generation features and Canvas (OpenAI’s AI-assisted workspace feature) are not supported1.
OpenAI states that “we recommend using it for challenging questions where reliability is more important than speed and where a few minutes of wait time is a worthwhile tradeoff”1, positioning o3-pro as a model specialized for professional tasks requiring high precision.
OpenAI’s o3-pro establishes a new benchmark in AI reasoning capabilities, outperforming competitors like Google’s Gemini 2.5 Pro and Anthropic’s Claude 4 Opus in critical areas such as mathematics and science. The model’s breakthrough performance - including 25.2% accuracy on Frontier Math where others achieve below 2% - demonstrates a significant leap in AI’s ability to tackle complex, specialized problems.
While the slower response times and current feature limitations present practical constraints, o3-pro’s strengths in reliability and accuracy make it valuable for high-stakes applications where precision matters more than speed. The integration of diverse tools from web search to Python execution, combined with enterprise-friendly API pricing, positions o3-pro as a powerful solution for organizations requiring advanced reasoning capabilities for research, technical analysis, and complex problem-solving tasks.
Sources
- OpenAI releases o3-pro, a souped-up version of its o3 AI reasoning model - TechCrunch (June 10, 2025)
- OpenAI o3 Released: Benchmarks and Comparison to o1 - Helicone (Detailed benchmarks of the o3 family)