The new GPT-6 Astra model utilizes looped transformers, where computation blocks are run multiple times with reused weights, effectively doubling the model's depth without increasing parameters. This architecture means the visible trace of the AI's reasoning is a post-hoc narration rather than a direct computation log. Consequently, evaluating AI agents and code by examining their 'chain of thought' becomes unreliable, as the generated text may not accurately reflect the actual computational process. Developers are advised to focus on observable behaviors like tool usage, file modifications, and final outputs for evaluation, rather than solely on the model's written reasoning. AI
IMPACT This architectural shift challenges current AI evaluation methods, requiring a move towards observable behaviors rather than solely relying on generated reasoning traces.
RANK_REASON New model release from a frontier lab with a novel architectural feature. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →