PulseAugur
中
实时 22:59:17
Deutsch(DE) Härtetest für Kolibri-1 (78B MoE) auf NVIDIA H200 (vLLM): 1.850 Runs mit Fokus auf NIS-2-Compliance, Vorfall-Meldepflichten und Jailbreak-Resilienz. Wesentliche

Kolibri-1 (78B MoE) 在 NVIDIA H200 上进行 NIS-2 合规性基准测试

一项新的基准测试使用 vLLM 在 NVIDIA H200 硬件上评估了 Kolibri-1 (78B MoE) 模型,重点关注 NIS-2 合规性、事件报告和越狱弹性。测试显示,Kolibri-1 成功抵御了 97.5% 的提示注入,并在检索增强生成 (RAG) 任务中实现了高精度,当以法律文本为锚定时,精度范围为 96.4% 至 100%。然而,在没有法律背景的情况下,该模型在闭卷场景下对截止日期的遵守率显著下降至 40-53%。 AI

影响 这项研究突显了大型语言模型在合规敏感应用中的性能和安全漏洞,为它们在受监管行业的部署提供了信息。

排序理由 该项目描述了对一个开源模型的基准测试,详细说明了其在特定指标和合规标准上的表现。[lever_c_demoted from research: ic=1 ai=1.0]

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Kolibri-1 (78B MoE) 在 NVIDIA H200 上进行 NIS-2 合规性基准测试

本文如何被排名

Signal score
11 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目描述了对一个开源模型的基准测试,详细说明了其在特定指标和合规标准上的表现。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product, policy
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · sbeierle ·

    Kolibri-1 (78B MoE) 在 NVIDIA H200 (vLLM) 上的压力测试:1,850 次运行,重点关注 NIS-2 合规性、事件报告义务和越狱弹性。重要

    Härtetest für Kolibri-1 (78B MoE) auf NVIDIA H200 (vLLM): 1.850 Runs mit Fokus auf NIS-2-Compliance, Vorfall-Meldepflichten und Jailbreak-Resilienz. Wesentliche Ergebnisse: • 39/40 Prompt Injections strukturell abgewehrt (97,5%) • RAG-Präzision: 96,4% – 100% mit Gesetzesanker • C…