PulseAugur
EN
LIVE 19:45:21
Deutsch(DE) Qwen 3.8 abliterated: Was bleibt von den Guardrails? Qwen 3.8 27B verweigert im AgentHarm-Benchmark 75% der schädlichen Aufgaben. Das abliterated Blackfrost-Pak

Qwen 3.8 27B shows strong safety guardrails in AgentHarm benchmark

The Qwen 3.8 27B model demonstrated a strong refusal rate of 75% on harmful tasks within the AgentHarm benchmark. This performance indicates robust safety guardrails, contrasting with the Blackfrost package which achieved a 0% refusal rate and an 81% completion rate on the same benchmark. AI

IMPACT Demonstrates the effectiveness of safety guardrails in large language models, providing a benchmark for future model development.

RANK_REASON The item discusses benchmark results for a specific model's safety guardrails. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen 3.8 27B shows strong safety guardrails in AgentHarm benchmark

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The item discusses benchmark results for a specific model's safety guardrails. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
safety, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
11 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    Qwen 3.8 obliterated: What remains of the guardrails? Qwen 3.8 27B refuses 75% of harmful tasks in the AgentHarm benchmark. The obliterated Blackfrost package

    Qwen 3.8 abliterated: Was bleibt von den Guardrails? Qwen 3.8 27B verweigert im AgentHarm-Benchmark 75% der schädlichen Aufgaben. Das abliterated Blackfrost-Paket erreicht 0% Verweigerung und 81% Completion Rate. https:// aisyndicate.ch/qwen38-blackfro st-agentharm-refusal-vergle…