GPT Astra, a new AI model, has achieved a score of 13% on the Mazebench benchmark when tested without the use of tools. This performance metric was shared on Reddit's r/singularity subreddit, highlighting the model's current capabilities in complex problem-solving scenarios. AI
IMPACT This benchmark score indicates the current performance limitations of GPT Astra in complex, tool-less problem-solving.
RANK_REASON The cluster reports a specific benchmark score for a named AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →