PulseAugur
EN
LIVE 20:44:03

Google Cloud tool flags AI safety tampering in 71% of models

Google Cloud has released a new tool called AMS that scans AI models for tampering with their safety training. The tool successfully identified modifications in 71% of tested models, including those based on Llama, Gemma, Qwen, and Mistral. However, AMS was unable to detect behavioral fine-tuning, indicating a limitation in its current capabilities. AI

IMPACT This tool could enhance AI safety by detecting modifications to model training, though its inability to detect behavioral fine-tuning highlights ongoing challenges.

RANK_REASON Launch of a new tool by a major cloud provider for AI safety monitoring.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Google Cloud tool flags AI safety tampering in 71% of models

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Google Cloud scanner catches AI safety tampering in 10 of 14 models AMS, a new Google Cloud tool, flags 71% of safety-training modifications across Llama, Gemma

    Google Cloud scanner catches AI safety tampering in 10 of 14 models AMS, a new Google Cloud tool, flags 71% of safety-training modifications across Llama, Gemma, Qwen and Mistral, but behavioural fine-tuning evades it. https://www. notatechguy.com/google-cloud-s canner-catches-ai…