OpenAI DevDay
OpenAI Ultrafast brings GPT-6 Astra to the API and Codex, with two different cost rules
The new speed tier is available for GPT‑6 Astra, but API token prices and Codex usage multipliers are different decisions. OpenAI’s published maximum‑speed figures also have different bases.
A speed tier on Astra, not another model
The API route exposes a request setting and separate rate limits
Codex and Work have plan gates and usage multipliers
For Codex, OpenAI’s Speed guide makes a narrower technical claim: GPT‑6 Astra Ultrafast can generate tokens up to eight times faster than Astra Standard in Codex. It explicitly separates token‑generation throughput from total task completion time. The same guide lists access for Pro $500 and eligible Enterprise and Edu plans, subject to the client, rollout and workspace settings. Enterprise access starts off by default; other self‑serve plans do not gain Ultrafast at launch merely by buying credits. A requirement for inference residency outside the United States also excludes a workspace.
The billing numbers sound similar to the speed numbers but describe something else. Astra Ultrafast draws down included subscription usage at eight times the Standard rate. Purchased credits and eligible Enterprise pay‑as‑you‑go usage are billed at six times Standard, subject to the Enterprise agreement. Codex and ChatGPT Work share usage. If Codex is authenticated with an API key, API token pricing applies instead of those ChatGPT credit multipliers. A buyer should check the route and payer before putting a single “6×” or “8×” cost estimate in a budget.
The practical test is cost per completed task
A useful evaluation keeps the model, prompt set, output length and tool sequence as comparable as possible across Standard and Ultrafast. Record time to first token, sustained token rate, full‑response time and every tool round trip. Then judge a complete task by its success criteria and total elapsed time. This design can reveal whether faster token streaming changes the actual user wait, or whether tool calls and external services dominate it. It is a proposed test, not one OpenTools has run.
Track API input, cached input, cache writes and output tokens separately against the published price schedule; for Codex/Work, use the plan’s documented allowance and credit rules. Compare cost per successful task, not just cost per token or the quickest demonstration. If quality changes or a request hits a regional or token‑rate gate, the simple speed multiple no longer answers the purchasing question. The launch is a credible reason to test a latency‑sensitive workload now, with its eligibility and accounting made explicit before rollout.
Related News
Sep 29, 2026
OpenAI Dots are always-on agents. Their most important launch feature is the control boundary
Dots combine a persistent cloud computer, 4,000-plus app connections and background work. OpenAI also draws a sharp line around approvals and read-only research.
Sep 29, 2026
OpenAI's Decisions API gives Luna a smaller job: choose from answers you define
The new API accepts text or images and returns a finite decision for classification, routing or an agent's next action. It is a limited preview, not a general release.
Sep 29, 2026
GPT-6.1 Sol launches at GPT-6 Sol prices, with cached input cut in half
OpenAI says GPT-6.1 Sol approaches Astra on agentic work without raising standard token prices. The launch data is promising, but its benchmark settings still matter.