A new analysis estimates that GPT-6 Astra's 50% time horizon (TH) for tasks without chain-of-thought prompting is likely between 8 minutes and 1 hour, with a probable median around 15-40 minutes. This estimate is based on performance on the "Think Fast" task suite, where Astra achieved high accuracy on many benchmarks, indicating near saturation. The study suggests that new, longer-horizon tasks are needed to accurately assess advanced models like Astra, as current benchmarks may not sufficiently challenge their capabilities. AI
IMPACT Suggests that current benchmarks may be insufficient for evaluating advanced AI models, highlighting the need for more challenging tasks.
RANK_REASON Analysis of an AI model's performance on a benchmark suite. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →