Tech Brew Ride Home · Thursday, August 13, 2026
OpenAI is introducing an 'Ultra Fast' service tier for its GPT-4.5 Sol model, capable of running up to 14 times faster than standard processing. This new tier, powered by Cerebrus, aims to generate up to 750 output tokens per second, making it suitable for latency-sensitive applications like live customer support and developer agents. Access is currently limited to a select group of customers for evaluation.
“OpenAI is previewing a new way to run its most capable GPT-4.5 model at dramatically higher speeds. The company says its new Ultra Fast service tier can run GPT-4.5 Sol up to 14 times faster than standard processing.”
“Ultra Fast mode launches first through the OpenAI API and is powered by Cerebrus. It can generate up to 750 output tokens per second, potentially bringing frontier-level performance to workflows where latency matters as much as model intelligence.”
“The company sees Ultra Fast supporting live or near production tasks, including voice customer support, commerce, developer agents, financial research, and security response.”