PulseAugur
中
实时 21:40:43
English(EN) I trained a model to be wrong 98% of the time and 96% sure about it. It took three tries.

开发者创建了一个AI模型,旨在高置信度地出错

一位开发者创建了一个名为Bev的AI模型,其设计目的是故意以高置信度提供错误答案。Bev在Qwen3.5-9B上进行了微调,旨在作为测试AI决策管道的对照案例。最初的训练尝试失败了,但通过使用现有的适配器并反转标签,最终成功,导致该模型大约98%的时间出错,同时对其错误响应保持96%的置信度。Bev可通过Hugging Face Spaces和Ollama获取,它既是一个幽默的“人工醉酒智能”,也是一个用于验证系统鲁棒性的工具。 AI

影响 通过充当旨在正确的模型对照案例,可作为独特的测试夹具,以验证AI决策管道的鲁棒性。

排序理由 发布了一个专门的、非前沿的AI模型用于测试目的。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

开发者创建了一个AI模型,旨在高置信度地出错

本文如何被排名

Signal score
5 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
发布了一个专门的、非前沿的AI模型用于测试目的。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/ricyoung ·

    我训练了一个模型,它98%的时间都出错,并且96%确信自己是错的。这花了三次尝试。

    <!-- SC_OFF --><div class="md"><p>Meet Bev.</p> <p>She is a decision model (the Jev / Nimble kind: you give her a situation and a question, she gives a probability for each answer), fine-tuned on Qwen3.5-9B to pick the worst answer on purpose.</p> <p>Try her in your browser: <a h…