PulseAugur
实时 09:23:23
English(EN) HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails

HoloAegis 框架通过几何推理提供零样本 LLM 安全性

研究人员推出了一种新颖的 LLM 安全护栏框架 HoloAegis,该框架利用冻结语义表征上的几何推理。这种方法避免了微调导致的表征失真与生成式裁判的高推理成本之间的权衡。HoloAegis 在多个基准测试中实现了最先进的准确性,具有最小的延迟且无需冷启动数据,并展示了强大的跨语言迁移能力。 AI

影响 这项研究提供了一种新的 LLM 安全方法,可以降低计算成本并提高准确性。

排序理由 该集群包含一篇学术论文,详细介绍了一种新的 LLM 安全方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

HoloAegis 框架通过几何推理提供零样本 LLM 安全性

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Tak Ho Alex Li, Kaijie Liu, Lik-Hang Lee, Kin Chung Ho, Ping Shum, Michael K. Ng ·

    HoloAegis:冻结表征、拓扑推断:用于零样本 LLM 护栏的最小参数化安全流形

    arXiv:2608.08485v1 Announce Type: new Abstract: Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inference costs. We challenge the prevailing paradigm by asking: can safety be achi…