PulseAugur
实时 15:45:26
English(EN) Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making

阿里巴巴Qwen发布AgentWorld语言模型用于环境模拟

阿里巴巴的Qwen团队推出了Qwen-AgentWorld,一个旨在模拟各种代理环境的新型语言世界模型。该模型侧重于训练大型语言模型理解和预测环境,而不仅仅是在其中行动。研究探索了两个主要途径:构建一个用于环境模拟的基础模型,以及研究世界建模如何增强代理训练,表明使用世界模型训练的代理可以优于在真实环境中训练的代理,并且预测性知识能有效地迁移到代理任务中。 AI

影响 这种方法可以通过提高代理对其运行环境的理解能力,从而实现更强大的代理,并可能加速复杂任务自动化的进展。

排序理由 前沿实验室模型发布,附带系统卡和基准测试结果。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 8 个来源。 我们如何撰写摘要 →

阿里巴巴Qwen发布AgentWorld语言模型用于环境模拟

报道来源 [8]

  1. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    🧠 Paradigm II — 智能体基础模型:世界建模作为智能体能力。

    🧠 Paradigm II — Agent Foundation Model: world modeling as agent capability. Single-turn, non-agentic environment prediction → tested directly on multi-turn, tool-calling agent tasks. No agentic RL, no task-specific tuning. Gains across 7 benchmarks, including 3 entirely https:/…

  2. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    第二部分:研究世界模型在智能体训练中的作用

    Part II: Investigating the Role of World Modeling in Agent Training 🔬 Paradigm I — Decoupled Simulation: world model as environment simulator for agent RL. The key is controllability: 1️⃣ Zero-shot generalization to 4k OOD OpenClaw environments → +4.3 Claw-Eval, +7.1 https://t…

  3. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    📊 AgentWorldBench:一个包含真实环境地面真实观测的7领域基准,由9个已建立基准上的5个前沿模型轨迹构建而成

    📊 AgentWorldBench: 7-domain benchmark with ground-truth observations from real environments, constructed from 5 frontier model trajectories on 9 established benchmarks. Results: Qwen-AgentWorld-397B-A17B achieves the highest overall score (58.71), outperforming Claude Opus 4.8 h…

  4. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    第一部分:构建环境模拟的基础模型

    Part I: Building the Foundation Model for Environment Simulation Faithful environment simulation requires multi-step causal reasoning, stateful tracking, and domain-specific knowledge. Frontier LLMs have some simulation ability from pretraining — but it's incidental, not an http…

  5. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    📣📣 认识 Qwen-AgentWorld — 一个原生语言世界模型,可在单一模式下模拟 7 种代理环境(MCP、Search、Terminal、SWE、Web、OS、Android)

    📣📣 Meet Qwen-AgentWorld — a native language world model that simulates 7 agent environments (MCP, Search, Terminal, SWE, Web, OS, Android) within a single model. Environment modeling is the training objective from day one, not a post-hoc adaptation. 🤔 LLMs are trained to be htt…

  6. arXiv cs.CL TIER_1 English(EN) · Guangfeng Cai, Kaibing Yang, Shuo He, Yu Li, Shengtian Yang, Jiaqi Lv, Lei Feng ·

    超越下一观测预测:用于序列决策的Agent自编世界模型

    arXiv:2606.25421v1 Announce Type: new Abstract: Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which…

  7. arXiv cs.CL TIER_1 English(EN) · Lei Feng ·

    超越下一观测预测:用于序列决策的Agent自编世界模型

    Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which may omit the dynamics most relevant to the agen…

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越下一观测预测:用于顺序决策的Agent自编世界模型

    Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which may omit the dynamics most relevant to the agen…