Limited preview · announced 29 Sep 2026

OpenAI Decisions API: what it is, what's real, and how to prepare

Everything publicly known about OpenAI's Decisions API — what it does, who can use it, reported speed, and pricing status — plus how to structure your code so switching later is trivial. Last reviewed 29 Sep 2026.

What the Decisions API is

OpenAI announced the Decisions API at DevDay on 29 Sep 2026. From OpenAI's recap:

“Decisions API enables real-time decision-making by focusing Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers. Developers supply context using text or images, and get back answers they can use to classify content, route requests, or choose an agent's next action.”

The idea is a dedicated call for a bounded choice instead of a general chat completion: you hand the model the context and the exact set of answers your code can act on, and it returns one of them — plus, per third-party reporting, a confidence score. Under the hood it is GPT-6 Luna, per The Decoder and OpenAI's developer account on X.

Status and timeline

  • 29 Sep 2026 — announced at OpenAI DevDay; “available in limited preview today with a broad release planned in the coming days” (OpenAI's words).
  • Today — no public endpoint, request schema, rate limits, or SLA published. Access is via the limited preview only.
  • Coming days — broad release planned per OpenAI. Dates beyond that are not published.

We update this page when OpenAI publishes docs, a schema, or pricing.

Speed — reported, not verified

OpenAI has not published latency or accuracy benchmarks for the Decisions API. Two third-party claims circulating:

  • The Decoder reports OpenAI claiming roughly 10× faster than asking Luna the same question via chat, and a computer-task simulation completing 76/78 steps correctly at ~230 ms per call.
  • byteiota reports ~150 ms typical latency, with the response returning one answer from your list plus a confidence score.

Treat both as reported claims until OpenAI publishes its own numbers.

Pricing

Not published. OpenAI has not announced Decisions API pricing. For reference only, GPT-6 Luna itself lists at $0.10 / 1M input tokens and $0.50 / 1M output tokens on OpenRouter's public register (via fourweekmba). Whether Decisions calls bill at Luna rates, a flat per-call fee, or something else is unknown. See our pricing watch page.

How to prepare your code now

You can't call OpenAI's endpoint yet, but you can ship the architecture it rewards today:

  1. Enumerate your answer space. Anywhere you currently parse free-text model output into an enum, write the enum down — it becomes your question's finite answer list.
  2. Separate deciding from acting. The model picks; your code checks thresholds and permissions. Keep a confidence floor (e.g. route to human review below 0.6).
  3. Abstract the call site. Put the decision call behind one function taking state + questions. Swapping providers later is a base-URL change.
  4. Instrument latency. If the 150–230 ms reports hold, decisions fit inside request paths. Measure what your current prompt-parse approach costs first.

Our playground and /v1/decisions API already run this exact request shape (state + finite-answer questions) on TypeSafe's Jev 1.13 — useful for prototyping the pattern while OpenAI's preview is closed.

Sources

Independent developer resource. Not affiliated with or endorsed by OpenAI or TypeSafe.