PulseAugur
实时 08:30:01
English(EN) Towards Multi-modal Multi-turn Safety: From Agentic Interaction to Strategic Alignment

新数据集和框架应对对话中的多模态大模型安全问题

研究人员推出 MINT-Safe,这是一个旨在解决多模态大语言模型(MLLMs)在扩展对话交互中安全问题的新数据集。该数据集包含 11,270 个多图像对话和 500 对拒绝式视觉问答(VQA)对,是通过多代理交互和文本到图像增强创建的。为了利用 MINT-Safe,该团队还开发了 TAD-Align 框架,该框架使用面向轮次的双目标奖励函数来动态识别和上调表现出不一致安全行为的对话轮次。在 Qwen2.5-VL-7B-InstructLLaVA-NeXT-7B 等模型上的实验显示,攻击成功率显著降低,无害性和有用性得到提高。 AI

影响 增强了对话环境中多模态大模型的安全协议,可能提高用户信任度并在敏感应用中得到部署。

排序理由 该集群包含一篇详细介绍多模态大模型新数据集和对齐框架的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新数据集和框架应对对话中的多模态大模型安全问题

本文如何被排名

Signal score
16 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍多模态大模型新数据集和对齐框架的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Han Zhu, Jiale Chen, Chengkun Cai, Shengjie Sun, Haoran Li, Yujin Zhou, Chi-Min Chan, Pengcheng Wen, Lei Li, Yike Guo, Sirui Han ·

    迈向多模态多轮安全:从代理交互到战略对齐

    arXiv:2601.04736v2 Announce Type: replace Abstract: Despite remarkable capability in multi-modal understanding, deploying Multi-modal Large Language Models (MLLMs) in open-ended conversational scenarios introduces safety risks that remain poorly addressed by existing alignment me…