English(EN)Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making
阿里巴巴Qwen发布AgentWorld语言模型用于环境模拟
作者PulseAugur 编辑部·[8 个来源]·
阿里巴巴的Qwen团队推出了Qwen-AgentWorld,一个旨在模拟各种代理环境的新型语言世界模型。该模型侧重于训练大型语言模型理解和预测环境,而不仅仅是在其中行动。研究探索了两个主要途径:构建一个用于环境模拟的基础模型,以及研究世界建模如何增强代理训练,表明使用世界模型训练的代理可以优于在真实环境中训练的代理,并且预测性知识能有效地迁移到代理任务中。
AI
🧠 Paradigm II — Agent Foundation Model: world modeling as agent capability.
Single-turn, non-agentic environment prediction → tested directly on multi-turn, tool-calling agent tasks. No agentic RL, no task-specific tuning.
Gains across 7 benchmarks, including 3 entirely https:/…
X — Qwen (Alibaba)
TIER_1English(EN)·Alibaba_Qwen·
Part II: Investigating the Role of World Modeling in Agent Training
🔬 Paradigm I — Decoupled Simulation: world model as environment simulator for agent RL.
The key is controllability:
1️⃣ Zero-shot generalization to 4k OOD OpenClaw environments → +4.3 Claw-Eval, +7.1 https://t…
X — Qwen (Alibaba)
TIER_1English(EN)·Alibaba_Qwen·
📊 AgentWorldBench: 7-domain benchmark with ground-truth observations from real environments, constructed from 5 frontier model trajectories on 9 established benchmarks.
Results: Qwen-AgentWorld-397B-A17B achieves the highest overall score (58.71), outperforming Claude Opus 4.8 h…
X — Qwen (Alibaba)
TIER_1English(EN)·Alibaba_Qwen·
Part I: Building the Foundation Model for Environment Simulation
Faithful environment simulation requires multi-step causal reasoning, stateful tracking, and domain-specific knowledge. Frontier LLMs have some simulation ability from pretraining — but it's incidental, not an http…
X — Qwen (Alibaba)
TIER_1English(EN)·Alibaba_Qwen·
📣📣 Meet Qwen-AgentWorld — a native language world model that simulates 7 agent environments (MCP, Search, Terminal, SWE, Web, OS, Android) within a single model. Environment modeling is the training objective from day one, not a post-hoc adaptation.
🤔 LLMs are trained to be htt…
arXiv:2606.25421v1 Announce Type: new Abstract: Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which…
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which may omit the dynamics most relevant to the agen…
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which may omit the dynamics most relevant to the agen…