Microsoft has launched a new in-house cybersecurity model, claiming it achieved a 95.95% score on the CyberGym benchmark at half the usual cost. However, this score was not reflected on the public leaderboard, and independent testers had not evaluated the model prior to its release. Furthermore, Microsoft's own documentation indicates the model performed with zero effectiveness in three specific exploit categories. AI
IMPACT Raises questions about the transparency and independent verification of AI models in critical security applications.
RANK_REASON The item describes the release of a new model by a major tech company, but it is not a frontier AI release and focuses on a specific application (cybersecurity) rather than general AI capabilities.
Read on Mastodon — sigmoid.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →