PulseAugur
实时 09:23:30
English(EN) Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts

新框架测试大型语言模型虚假信息检测器对抗对抗性攻击的能力

研究人员开发了一个名为 Build it, Break it, Repeat (BiBiR) 的新框架,用于测试虚假信息检测模型对抗由大型语言模型操纵的内容的鲁棒性。这种迭代方法模拟了系统性地修改虚假信息帖子以规避分类的对抗性条件。在实验中,最佳的对抗性转换包括回译和基于大型语言模型角色重写,在保持原意的情况下实现了 95% 的标签翻转率。表现最佳的检测模型,一个具有动态锚点切换 (DASS) 的三元组对比架构,在面对这些复杂的攻击时达到了 72.68% 的准确率,显著优于经过微调的 e5-small-LoRA 基线。 AI

影响 这项研究强调了对人工智能驱动的虚假信息检测需要更鲁棒的评估方法,这对于维护平台完整性至关重要。

排序理由 学术论文,详细介绍了评估人工智能模型鲁棒性的新方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新框架测试大型语言模型虚假信息检测器对抗对抗性攻击的能力

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Kevin Thomas, Milosz Kasprzyk, Reuel C Igbokwe Onuigbo, Elliott Pert, Cameron Tovey, Jo\~ao A. Leite, Olesya Razuvayevskaya, Carolina Scarton ·

    构建、破坏、重复:社交媒体帖子中LLM操纵的虚假信息检测基准测试与改进

    arXiv:2608.09510v1 Announce Type: cross Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measuring detec…