PulseAugur
实时 14:45:37

新基准 MulRobBench 测试无人机代理的安全和安保策略决策

研究人员推出 MulRobBench,这是一个旨在评估智能城市环境中多模态无人机 (UAV) 代理的新基准。该基准侧重于决策,整合了真实的无人机观测、安全策略和网络物理安全约束。MulRobBench 旨在评估这些代理在退化条件和模糊指令下的运行能力,将语义评分与策略合规性和不安全行为识别等结构性诊断相结合。初步评估显示,即使是表现最好的多模态模型,语义协议-决策得分也仅为 0.5141,凸显了在实现自主无人机可信决策方面存在的重大挑战。 AI

影响 该基准有望推动复杂城市环境中自主无人机系统安全性和安全性合规性的改进。

排序理由 该条目描述了一个用于评估 AI 代理的新基准,发布在 arXiv 上。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.MA (Multiagent) 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新基准 MulRobBench 测试无人机代理的安全和安保策略决策

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Belal S. Alsinglawi, Weizheng Wang, Junyi Wu, Yi Jiang, Lianhai Lin, Merouane Debbah, Izzat Alsmadi ·

    MulRobBench:安全且符合安全策略的多模态无人机代理的决策级基准测试

    arXiv:2607.23870v1 Announce Type: cross Abstract: Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing…

  2. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Izzat Alsmadi ·

    MulRobBench:安全且符合安全策略的多模态无人机代理的决策级基准测试

    Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing UAV and multimodal benchmarks evaluate perception…