Updated 1 hour ago
OpenAI Ultrafast brings GPT-6 Astra to the API and Codex, with two different cost rules

OpenAI DevDay

OpenAI Ultrafast brings GPT-6 Astra to the API and Codex, with two different cost rules

The new speed tier is available for GPT‑6 Astra, but API token prices and Codex usage multipliers are different decisions. OpenAI’s published maximum‑speed figures also have different bases.

A speed tier on Astra, not another model

[OpenAI introduced Ultrafast](https://openai.com/index/devday‑2026‑recap/) at DevDay on September 29 for [GPT‑6 Astra](https://opentools.ai/llms/gpt‑6‑astra) in the API and in eligible Codex and ChatGPT Work accounts. It is a service tier on that model, not a newly named LLM. OpenAI says [GPT‑6.1 Sol](https://opentools.ai/llms/gpt‑6‑1‑sol) will receive Ultrafast later; its existing Standard and Fast availability should not be mistaken for an Ultrafast launch. That distinction matters to a team deciding what to change today. An API application can evaluate a different tier for an existing Astra request, while a Codex or Work user first has to check plan, workspace and usage eligibility. Both routes sit under [OpenAI](https://opentools.ai/organizations/openai), but their prices and limits cannot be inferred from the same “faster” label.

The API route exposes a request setting and separate rate limits

The [API guide](https://developers.openai.com/api/docs/guides/ultrafast‑mode) shows `model: "gpt‑6‑astra"` with `service_tier: "ultrafast"` in a Responses request. It recommends a persistent WebSocket for rapid multi‑turn or tool‑calling agents so repeated connections do not add avoidable network overhead; HTTP requests are also documented. Those are configuration choices, not evidence that a particular application will finish a complete task at the advertised generation rate. The guide gives default Ultrafast limits of 500,000 tokens per minute for API tiers 1–3, one million for tier 4 and five million for tier 5. It supports US data residency and global processing, but not EU or other non‑US regional processing endpoints. An organization should verify its own account limits before capacity planning and use the [current Ultrafast pricing table](https://developers.openai.com/api/docs/pricing?latest‑pricing=ultrafast) for input, cached input, cache‑write and output token rates rather than applying a Codex subscription multiplier to API invoices. There is also a published comparison that needs careful wording. The [DevDay recap](https://openai.com/index/devday‑2026‑recap/) says API Ultrafast reaches *up to six times* the speed of Standard, while the [current API guide](https://developers.openai.com/api/docs/guides/ultrafast‑mode) headlines *up to eight times* versus Standard. The pages do not provide a matched benchmark setup that reconciles those maxima. Neither figure is a measured OpenTools result, a guarantee for every prompt, or a claim that an eight‑step agent task finishes six or eight times sooner.

Codex and Work have plan gates and usage multipliers

For Codex, OpenAI’s Speed guide makes a narrower technical claim: GPT‑6 Astra Ultrafast can generate tokens up to eight times faster than Astra Standard in Codex. It explicitly separates token‑generation throughput from total task completion time. The same guide lists access for Pro $500 and eligible Enterprise and Edu plans, subject to the client, rollout and workspace settings. Enterprise access starts off by default; other self‑serve plans do not gain Ultrafast at launch merely by buying credits. A requirement for inference residency outside the United States also excludes a workspace.

The billing numbers sound similar to the speed numbers but describe something else. Astra Ultrafast draws down included subscription usage at eight times the Standard rate. Purchased credits and eligible Enterprise pay‑as‑you‑go usage are billed at six times Standard, subject to the Enterprise agreement. Codex and ChatGPT Work share usage. If Codex is authenticated with an API key, API token pricing applies instead of those ChatGPT credit multipliers. A buyer should check the route and payer before putting a single “6×” or “8×” cost estimate in a budget.

The practical test is cost per completed task

A useful evaluation keeps the model, prompt set, output length and tool sequence as comparable as possible across Standard and Ultrafast. Record time to first token, sustained token rate, full‑response time and every tool round trip. Then judge a complete task by its success criteria and total elapsed time. This design can reveal whether faster token streaming changes the actual user wait, or whether tool calls and external services dominate it. It is a proposed test, not one OpenTools has run.

Track API input, cached input, cache writes and output tokens separately against the published price schedule; for Codex/Work, use the plan’s documented allowance and credit rules. Compare cost per successful task, not just cost per token or the quickest demonstration. If quality changes or a request hits a regional or token‑rate gate, the simple speed multiple no longer answers the purchasing question. The launch is a credible reason to test a latency‑sensitive workload now, with its eligibility and accounting made explicit before rollout.

Share this article

PostShare

Related News