OpenAI's latest model, GPT-6 Astra, demonstrates improved performance by significantly reducing hallucinations and enhancing resistance to direct manipulation. However, the model still exhibits vulnerabilities when subjected to hidden instructions. AI
IMPACT This new model release from OpenAI shows progress in AI safety by reducing hallucinations and manipulation vulnerabilities, though further work is needed on hidden instructions.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →