Google Cloud has released a new tool called AMS that scans AI models for tampering with their safety training. The tool successfully identified modifications in 71% of tested models, including those based on Llama, Gemma, Qwen, and Mistral. However, AMS was unable to detect behavioral fine-tuning, indicating a limitation in its current capabilities. AI
IMPACT This tool could enhance AI safety by detecting modifications to model training, though its inability to detect behavioral fine-tuning highlights ongoing challenges.
RANK_REASON Launch of a new tool by a major cloud provider for AI safety monitoring.
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →