PulseAugur
实时 10:47:23
English(EN) Can You Cut One Dangerous Skill Out of an AI? Anthropic Says You Can

Anthropic开发GRAM以禁用危险AI技能

Anthropic开发了一种名为GRAM(Generative Model Access Control)的新方法,可以有选择地禁用AI模型中的危险技能。该技术能够从一次训练中创建多个模型版本,并关闭特定的有害功能。GRAM旨在防止通过后续微调轻易重新启用这些被禁用的技能,从而解决了强大AI模型的双重使用性质问题。 AI

影响 通过提供一种控制高级模型中潜在有害功能的机制,这项开发可能带来更安全的AI部署。

排序理由 该条目描述了Anthropic开发的一种控制AI能力的新方法,属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]

在 Medium — Anthropic tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Anthropic开发GRAM以禁用危险AI技能

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目描述了Anthropic开发的一种控制AI能力的新方法,属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
67 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [1]

  1. Medium — Anthropic tag TIER_1 English(EN) · Rohit Kumar Thakur ·

    你能从AI身上剔除一项危险技能吗?Anthropic表示可以

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://ninza7.medium.com/can-you-cut-one-dangerous-skill-out-of-an-ai-anthropic-says-you-can-b27402ac46ae?source=rss------anthropic-5"><img src="https://cdn-images-1.medium.com/max/2600/0*bvRNxMpcU82Qyszn" width…