OpenAI has introduced a new API service tier called Ultrafast Mode, designed to significantly accelerate the performance of its GPT-5.6 "Sol" model. This new tier can process requests up to 14 times faster than previous offerings, achieving an output rate of up to 750 tokens per second. The enhanced speed is made possible by leveraging Cerebras hardware. AI
IMPACT Accelerates inference for GPT-5.6 "Sol", potentially lowering costs and enabling new real-time applications.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →