Summary

OpenAI previewed Ultrafast, a new API service tier that runs GPT-5.6 Sol at up to roughly 750 output tokens per second, up to 14x faster than standard processing, using Cerebras wafer-scale hardware. It launches first in the API to a limited group of enterprise customers.

What changed

OpenAI introduced an Ultrafast API service tier for GPT-5.6 Sol powered by Cerebras wafer-scale chips, delivering up to about 750 tokens per second (up to 14x standard speed), available in limited preview to select enterprise customers, priced above the existing Fast tier with exact pricing undisclosed.

Why it matters

Real-time latency usually forces teams to drop down to smaller models. Putting a frontier model into a low-latency tier lets latency-sensitive agentic products in coding, commerce, customer support, and financial research keep top-tier reasoning without the speed tradeoff, and it signals deepening OpenAI-Cerebras inference dependence.

Evidence excerpt

Ultrafast runs GPT-5.6 Sol at up to 750 output tokens per second and up to 14x faster than Standard processing, powered by Cerebras wafer-scale hardware.

Sources