OpenAI Releases the Decisions API in Beta, Returning Only Typed Answers - Probability, Choice, and Score, Billed on Input Alone
On October 6, 2026, OpenAI released the Decisions API in beta. It evaluates text or images and returns only typed answers. gpt-6-luna is the sole model, answers come in three forms (predicate, choice, score), and pricing is $0.10 per million input tokens with no charge for output or caching.
On October 6, 2026, OpenAI put the Decisions API into beta2. It does not write prose. You hand it text or images along with a set of questions, and it replies in one of three structured forms: the likelihood that some condition holds, one pick from options you defined, or a position on a graded scale. Requests go to a dedicated POST /v1/decisions endpoint, and for now gpt-6-luna is the only model on offer1.
OpenAI says the API answers roughly ten times faster than the Responses API, but neither the guide nor the changelog describes how that was measured. You are billed $0.10 per million tokens, and only for what you send in. It is a public beta for now, and OpenAI expects general availability within weeks.
Three answer types, and nothing else
Each request is built from three pieces: the evaluating model (model), the evidence (input), and what you want judged (questions). Each question declares a type, and the type decides what comes back1.
- predicate: whether a condition is true, returned as a value between 0 and 1. The guide’s example asks whether a product photo shows physical damage such as a crack or dent
- choice: one value out of options you supply. The example routes a “charged twice” complaint to billing, technical, shipping, or other
- score: a rating on ordered levels. Because the result weights each level’s zero-based index by its probability, it can land between levels. The guide rates bug severity on three levels, where a 0.1 / 0.7 / 0.2 split gives 1.1
Choice and score answers also carry a per-option probability array and a separate confidence value. For deciding when to hand an item to a person, the guide recommends tuning cutoffs on labeled data from your own application, weighing what a false alarm costs against what a miss costs. For choice questions whose options might not cover every input, it suggests adding a catch-all such as other.
Unrelated questions can share one request, and each can use its own type. When a later question depends on an earlier answer, they have to go in separate calls.
Input-only pricing, base64-only images
Billing works differently from ordinary model calls. Using gpt-6-luna through /v1/decisions costs $0.10 for every million tokens of input; OpenAI lists no charges for output tokens or for cache reads and writes1. Surcharges for regional processing and multipliers for long-context input still apply, and any gpt-6-luna call outside this endpoint follows the regular price list.
When GPT-6 Luna launched on September 22, its standard rates put input at $0.10, cached input at $0.01, and output at $0.502. The input figure is unchanged; what the Decisions API drops is the separate output and cache line items. With this endpoint, the size of the bill comes down to how many input tokens you send.
There are limits. Images must be embedded as base64 data URLs; the endpoint rejects HTTP/HTTPS image links and file_id references1. A pipeline that currently points the model at object-storage URLs will need changes before it can move over.
On data handling, eligible customers can run it for HIPAA workloads and with ZDR (Zero Data Retention), and both the US and Europe (EEA plus Switzerland) are supported for data residency and in-region processing. For organizations that have to settle retention policy and processing location up front, having that scope stated during the beta gives them something concrete to evaluate.
Decision models and Strands Decider 2B
TechCrunch reports that models emitting outcome probabilities instead of text have been a subject of attention in the AI industry since TypeSafe AI released Jev in September, with OpenAI and Amazon following with rival models3. On the same October 6, Musubi announced PolicyLM-1.7B, an open-weight decision model for content moderation. OpenAI’s own documentation, for its part, does not use the term “decision model” for the Decisions API.
On October 4 we covered Strands Decider 2B, which strands-labs released on October 1. It also answers in three forms, a pick, a yes/no, and a score, close to the Decisions API’s three types. Where they differ is location. Strands Decider 2B is a 1.9-billion-parameter model with weights under Apache 2.0 that you run on your own GPU or CPU. The Decisions API calls gpt-6-luna in OpenAI’s cloud, with no choice of model for now.
How each handles confidence is also worth comparing. strands-labs pointed out that a decision model returns a confidence for each judgment, and said frontier LLM inference APIs do not provide it. The Decisions API returns a probability for predicates, and a probability distribution plus confidence for choice and score1. Which route fits will likely depend on whether your data can leave your environment, whether you can operate GPUs, and how many judgments you need to run.
Where Structured Outputs still fits
The documentation also lays out how this relates to other features. If one of the three answer types is enough, use the Decisions API. To generate an object shaped by a JSON schema you define, for instance pulled-out fields or a written rationale, use Structured Outputs in the Responses API. To have a model request a tool with arguments, use function calling1.
Places where Structured Outputs currently returns a single field like {"category": "billing"} look like candidates for moving over. Cases that need a written rationale, or several fields pulled out at once, sit outside what the Decisions API does. The guide closes by pointing to voice use: paired with client delegation in the Live API, Decisions can pick which action to take from a spoken request.
During the beta there is one supported model, and OpenAI’s only statement on speed is the “about 10x” figure with no test conditions attached. If you build classification or routing on LLM calls, the sensible first step is to run your own inputs through the Playground and compare speed, accuracy, and cost against your current setup. Whether pricing or model support changes at general availability is another thing to watch.
Sources
- Decisions - OpenAI API documentation (Decisions API guide)
- Changelog - OpenAI API changelog (Decisions API release on October 6, 2026; GPT-6 Luna release on September 22)
- How AI decision models could change content moderation - TechCrunch (October 6, 2026)
Was this article helpful?
Thank you!
Received. Thank you!