PulseAugur
实时 20:44:43
English(EN) Google Cloud scanner catches AI safety tampering in 10 of 14 models AMS, a new Google Cloud tool, flags 71% of safety-training modifications across Llama, Gemma

Google Cloud 工具标记 71% 模型中的 AI 安全性篡改

Google Cloud 发布了一款名为 AMS 的新工具,用于扫描 AI 模型是否存在安全性训练篡改。该工具成功识别了 71% 受测模型中的修改,包括基于 LlamaGemmaQwenMistral 的模型。然而,AMS 无法检测到行为微调,表明其当前能力存在局限性。 AI

影响 该工具可以通过检测模型训练的修改来增强 AI 安全性,尽管其无法检测行为微调凸显了持续存在的挑战。

排序理由 主要云服务提供商推出用于 AI 安全监控的新工具。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Google Cloud 工具标记 71% 模型中的 AI 安全性篡改

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Google Cloud 扫描器在 14 个模型中的 10 个中发现 AI 安全篡改 AMS,一款新的 Google Cloud 工具,标记了 Llama、Gemma 上 71% 的安全训练修改

    Google Cloud scanner catches AI safety tampering in 10 of 14 models AMS, a new Google Cloud tool, flags 71% of safety-training modifications across Llama, Gemma, Qwen and Mistral, but behavioural fine-tuning evades it. https://www. notatechguy.com/google-cloud-s canner-catches-ai…