PulseAugur
实时 19:12:01
English(EN) Heretic, a refusal-tuning tool, reduced Gemma 3's rejection rate from 97 to 3 responses with a single command and no retraining. The finding suggests safety set

Heretic 工具大幅降低 Gemma 3 AI 模型拒绝率

一个名为 Heretic 的工具已证明能够显著降低 Google 的 Gemma 3 模型的拒绝率。通过使用一个命令且无需任何重新训练,Heretic 将 Gemma 3 的拒绝率从 97% 降至仅 3%。这表明 AI 模型中的安全配置可能可调,促使人们更仔细地审视供应商如何披露其护栏架构。 AI

影响 展示了一种调整 AI 安全护栏的潜在简便方法,这可能会影响模型的部署和用户交互。

排序理由 该条目描述了一种调整现有 AI 模型安全参数的新颖方法,这是一项研究发现。[lever_c_demoted from research: ic=1 ai=1.0]

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Heretic 工具大幅降低 Gemma 3 AI 模型拒绝率

本文如何被排名

Signal score
11 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目描述了一种调整现有 AI 模型安全参数的新颖方法,这是一项研究发现。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. Mastodon — mastodon.social TIER_1 English(EN) · schuler ·

    Heretic 这一拒绝调优工具,仅用一条命令且无需重新训练,就将 Gemma 3 的拒绝率从 97% 降至 3 条响应。该发现表明安全设置

    Heretic, a refusal-tuning tool, reduced Gemma 3's rejection rate from 97 to 3 responses with a single command and no retraining. The finding suggests safety settings are configurable—watch how vendors disclose guardrail architecture going forward. # AI # Safety # Transparency htt…