OpenAI Previews Ultrafast GPT-5.6 Sol API Tier

OpenAI previewed Ultrafast on August 13, a limited-access API tier that runs GPT-5.6 Sol at up to 750 output tokens per second, or up to 14x Standard processing speed. OpenAI and Cerebras report that Cerebras hardware powers the tier, which is initially available to a select group of customers for latency-sensitive workloads including voice, support, incident response, commerce, finance, and security.
OpenAI previewed Ultrafast, a limited-access API service tier for GPT-5.6 Sol, on August 13. According to OpenAI, the tier runs the model at up to 750 output tokens per second, or up to 14x the speed of Standard processing, and is powered by Cerebras hardware.
The release is a serving option rather than a new foundation model. OpenAI states that Ultrafast launches first through its API and is being tested with an initial group of customers. Its signup page says capacity is limited and that customer inclusion will be evaluated based on workload fit and availability.
A latency-focused tier
OpenAI identifies incident response, reliability work, financial research and security, customer support and voice, commerce, and live research and experimentation as candidate use cases. In its announcement, the company argues that prior real-time deployments commonly required smaller or more specialized models, while Ultrafast is intended to provide the capabilities of GPT-5.6 Sol at lower latency.
Sachin Katti, OpenAI's VP of Compute Strategy and GPT-Infra, described the preview in a Cerebras announcement as an effort to learn where reduced latency creates meaningful customer value. "We're starting with a small group of customers to learn where that speed creates meaningful value, and we'll use those learnings to inform how we expand the service over time," Katti said.
Cerebras describes its role as powering the Ultrafast inference service. The company claims that the tier provides the same intelligence as Standard processing, though OpenAI's public announcement primarily specifies the throughput improvement rather than publishing independent quality measurements for the new serving configuration.
Vendor benchmarks require context
Cerebras also published comparative results that should be read as vendor-reported benchmarks. It states that, in its Humanity's Last Exam evaluation, GPT-5.6 Sol Ultrafast answered 2,500 questions in 11 hours and 11 minutes, compared with 78 hours and 27 minutes for Anthropic's Claude Fable 5. Cerebras says the tests used GPT-5.6 Sol Ultrafast with Codex on xhigh reasoning on July 10, and Claude Fable 5 with Claude Code on xhigh reasoning from July 13 to 15.
The company further reports a 5.6x end-to-end speedup on GDP-Val without quality degradation, and cites Artificial Analysis output-speed data to claim a 5x advantage over Claude Opus 4.8 Fast mode and an 11x advantage over Claude Fable 5. Those comparisons are useful indicators of the throughput target, but their applicability depends on model settings, tool use, prompting, output length, concurrency, and the benchmark harness.
For ML teams building interactive systems, output-token throughput is only one part of observed latency. End-to-end response time also includes request routing, prompt ingestion, retrieval, tool calls, safety checks, application-side orchestration, and streaming behavior. Companies deploying comparable high-speed inference tiers often need to measure these stages separately, particularly for voice agents, incident workflows, and multi-step tool-using systems where each model turn can add delay.
Implications for inference infrastructure
The announcement gives Cerebras a high-profile production inference deployment for an OpenAI API tier. Cerebras CEO Andrew Feldman characterized the collaboration as evidence that speed and frontier-model capability need not be mutually exclusive, while OpenAI's materials frame the preview around workloads where delays affect usability.
The limited preview leaves several practical details undisclosed in the published materials, including pricing, regional availability, rate limits, context-window behavior, and formal service-level commitments. Teams considering the tier will need those operating details, alongside workload-specific latency and quality tests, before treating the headline 750-token-per-second figure as an application-level performance guarantee.
Key Points
- 1OpenAI's limited Ultrafast preview targets up to 750 output tokens per second, making inference latency a first-class API selection criterion.
- 2Cerebras-reported benchmark results suggest substantial throughput gains, but practitioners need independent workload tests across tools, prompts, concurrency, and end-to-end latency.
- 3For real-time agents, faster generation can reduce per-turn delay, while retrieval, routing, and tool execution remain material sources of application latency.
Scoring Rationale
This is a notable production inference collaboration involving OpenAI, Cerebras, and a reported 14x throughput increase for a frontier-model API tier. It is particularly relevant to teams building latency-sensitive agents and real-time applications, although access remains limited and key operational details are not yet public.
Sources
Primary source and supporting public references used for this report.
View 4 more sources
- Accelerating GPT-5.6 Sol Ultrafast with OpenAIcerebras.ai
- Cerebras Powers Ultrafast Mode for OpenAI's GPT-5.6 Solinvestors.cerebras.ai
- OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speedtechcrunch.com
- OpenAI’s new Ultrafast mode runs GPT-5.6 Sol 14 times faster, on Cerebras chipsthenextweb.com
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems
