OpenAI previews Ultrafast API tier for GPT-5.6 Sol

OpenAI's Ultrafast API tier is in limited preview for select customers, with broader access planned as capacity expands.

· 2 min read
OpenAI

OpenAI has opened a limited preview of Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol at up to 14 times the speed of Standard processing. Powered by Cerebras, the mode can generate up to 750 output tokens per second. Access is initially restricted to a select group of customers, with a wider rollout planned as capacity grows.

The service is designed for products and workflows where delays can determine whether an answer is still useful. OpenAI says achieving real-time speeds has often required teams to choose a smaller or specialized model. Ultrafast instead puts the company's most intelligent model into low-latency settings, seeking to deliver more useful work per second without making that tradeoff.

The immediate targets span incident response, finance, security, customer support, voice, commerce, and research. Teams could analyze logs, code changes, transactions, or market signals while events are unfolding. Voice and support systems could resolve multi-step requests without breaking a conversation, while commerce tools could check inventory, tailor recommendations, and address checkout problems before a shopper leaves. Researchers could also test and adjust work in shorter cycles.

OpenAI is already using the tier internally. During incidents, its developers have applied it to logs, traces, team conversations, follow-up checks, and preparation or validation of fixes, while engineers retain responsibility for judgment and deployment. Research teams are using it across connected tools to search knowledge sources, query data, and organize findings. OpenAI says some experiment loops that once ran overnight can now support several iterations within a workday.

Early testing includes Jane Street, Podium, Basis, and Rogo. Their feedback centers on more focused coding sessions, faster complex voice calls, low-latency applications built around a frontier model, and financial research that feels closer to a live exchange.

Ultrafast extends OpenAI's partnership with Cerebras, whose infrastructure supports GPT-5.6 Sol at the stated output rate. The preview is available through the API, and OpenAI is collecting input from initial customers to guide the service as capacity expands. A sign-up form is available for access updates, but the company has not announced pricing or a general-availability timeline.

Source