OpenAI has introduced Ultrafast, a new service tier designed to dramatically increase the speed of its GPT-5.6 Sol model. Powered by Cerebras, Ultrafast can generate up to 750 output tokens per second, with OpenAI claiming speeds of up to 14× faster than Standard processing.

Ultrafast is launching initially through the OpenAI API for a select group of customers. The company says the service is designed for latency-sensitive workloads such as developer agents, real-time voice applications, financial research, customer support, and security response.

Why It Matters

Faster AI inference could make autonomous agents and real-time AI applications significantly more responsive. For businesses, reducing model response times could also accelerate workflows where every second matters, particularly in software development, cybersecurity, and customer operations.

Conclusion

The GPT-5.6 Sol Ultrafast launch highlights a growing industry focus on AI inference speed alongside model intelligence. With Cerebras powering the new service tier, OpenAI is pushing toward faster AI workloads for production and near-production applications.