A new research paper introduces StateM, an agent-native runtime designed to improve the performance of long-horizon AI agents. StateM focuses on enhancing the execution system around an agent without altering its core model weights, by organizing execution around durable states, phase-local context, and versioned procedural practices. The system demonstrated significant accuracy improvements on the Terminal-Bench 2.1 benchmark, boosting GPT-5.5 xhigh to 92.1% and achieving 95.3% with GPT-5.6 Sol xhigh. Furthermore, StateM improved DeepSeek-V4 Flash's accuracy from 82.7% to 88.1% with minimal adaptation cost and reduced API usage significantly. AI
IMPACT Enhances long-horizon agent capabilities and reduces inference costs, potentially accelerating adoption of complex AI workflows.
RANK_REASON Research paper detailing a new methodology for improving AI agent performance. [lever_c_demoted from research: ic=1 ai=1.0]
- BusinessBench
- DeepSeek-V4 Flash
- GPT-5.5 xhigh
- GPT-5.6
- GPT-5.6 Luna
- GPT-5.6 Sol
- GPT-5.6 Sol Ultra
- GPT-5.6 Sol xhigh
- Terminal-Bench 2.1
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →