OpenAI Releases the Agents API in Public Beta - Borrowing the Codex Harness as a Managed Service
OpenAI released the Agents API in public beta on September 10, 2026. The Codex harness comes as a managed service, with orchestration, context compaction and recovery run on OpenAI's side, and with US-only data residency and no ZDR support written into the docs.
On September 10, 2026, OpenAI announced in its API changelog that the Agents API is now in public beta. The Codex harness is offered as a managed service: per the changelog, the work of orchestrating sessions, compacting context, and recovering from failures moves to OpenAI’s side1.
For anyone wiring agents into an in-house system, the question this raises is who holds the loop and the session state. OpenAI’s documentation lines up three agent runtimes — the Agents API, the Agents SDK, and the Responses API — and only the Agents API sits on the side where “OpenAI runs a managed Codex harness”3.
Create a session, hand it a task
The official documentation describes the Agents API as giving your application access to the Codex harness through an OpenAI-managed API2. The split of responsibilities is stated plainly. Everything to do with keeping a session alive — orchestration, context compaction, recovery — belongs to OpenAI; supplying the tools and picking where the agent runs belongs to the application.
Four concepts organize the surface. An agent is the model, instructions, tools, and MCP servers available to it. An environment is the optional sandbox or computer in which the agent reads files, picks up skills, and issues commands. A session is a durable instance of an agent that works on tasks and responds to input. Events and items are the inputs sent in and the output produced during a session2. The usage flow runs in four steps: create a session, give it a task, follow progress by streaming or webhooks, then send another task to the same session or steer the agent mid-turn.
Seven capabilities are listed for the managed harness2: it runs commands and code inside a sandbox, pulls in the skills and instructions that apply, reaches outside data through tools or MCP, accepts steering mid-task, compresses earlier work into summaries so the context window holds, splits a job into subtasks for subagents, and picks a session back up after an interruption. If you have ever built agent infrastructure yourself, that list maps almost exactly onto the parts that are tedious to write. Being able to attach MCP servers as tools will look familiar to readers who followed the large MCP specification revision finalized in July.
The sample code specifies GPT-6 Astra as the model, sets a cap on concurrently running subagents through multi_agent, and lists an MCP server and web search as tools2. The endpoint is POST /v1/agents/sessions, and requests carry an OpenAI-Beta: agents=v1 header.
Three execution environments, billing that stacks up
The execution environment can be an OpenAI-hosted sandbox, a sandbox on your own infrastructure or from a supported provider, or no sandbox at all3. Inside a sandbox, the agent can execute code, edit files, connect to MCP servers, and produce artifacts.
Billing stacks: model usage is charged at the selected model’s API rates, OpenAI tools at their standard rates, and OpenAI-hosted sandboxes at standard container rates2. The documentation carries no mention of a separate fee for the Agents API itself. Choosing a self-hosted execution environment would presumably shift the container-rate portion onto your own infrastructure costs. Keeping agent execution inside your own network is the same shape as Cursor’s self-hosted machines.
The documentation links five worked examples, among them incident response, data analysis, and GitHub issue investigation2. They lean toward work that does not finish in a single response and that reaches for several tools before a result comes back.
Integration effort weighed against what you hand over
OpenAI compares the three runtimes side by side in its documentation3. The Agents API is for “long-running tasks where OpenAI manages the agent and saves its progress,” with integration effort listed as Low. The Agents SDK is for “building agents with custom tools and workflows in your application,” at Medium. The Responses API is for “calling models directly or building an agent from scratch,” at High. State between tasks splits the same way: saved session configuration, turns, and items for the Agents API; your own storage or SDK sessions for the Agents SDK; manual history for the Responses API.
Lower effort comes with something handed over. The documentation states that the Agents API currently supports data residency only in the United States and does not support Zero Data Retention (ZDR), and adds that choosing a self-hosted sandbox does not make the Agents API ZDR-eligible2. Session state is retained so that conversation context does not have to be rebuilt across turns, and sessions and published artifacts can be deleted when they are no longer needed.
That constraint is where the choice tends to split. Where data residency or ZDR is part of a procurement requirement, adopting the Agents API as-is will be difficult, and keeping your own loop on the Agents SDK or the Responses API fits the requirement better. Where the requirements are looser — an internal tool, say — handing off harness maintenance entirely has real value. OpenAI also offers Presence as an enterprise agent operations platform, but the Agents API is not a finished operations product; it is supplied as a part to build into your own application.
The documentation also cautions that an Agents API session, an SDK session, a Responses conversation, and a sandbox are different resources, and that you should follow the state and cleanup instructions for whichever runtime you choose3. With three runtimes coexisting, cleaning up after a mixed deployment falls to the user.
The public-beta label is worth factoring in as well. The changelog gives no timing for general availability1, and the surface sits under a beta namespace in the SDKs and behind an OpenAI-Beta header over HTTP. Given that the CLI carrying the same Codex name also ships updates at short intervals, it is safer to treat the API’s shape as still moving for now.
Sources
- Changelog - OpenAI API - OpenAI official API changelog (September 10, 2026 entry)
- Agents API - OpenAI official documentation
- Agents - OpenAI official documentation (agent runtime comparison)
Was this article helpful?
Thank you!
Received. Thank you!