PulseAugur
中
实时 19:54:19
English(EN) Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making

阿里巴巴Qwen发布AgentWorld语言模型用于环境模拟

阿里巴巴的Qwen团队推出了Qwen-AgentWorld,一个旨在模拟各种代理环境的新型语言世界模型。该模型侧重于训练大型语言模型理解和预测环境,而不仅仅是在其中行动。研究探索了两个主要途径:构建一个用于环境模拟的基础模型,以及研究世界建模如何增强代理训练,表明使用世界模型训练的代理可以优于在真实环境中训练的代理,并且预测性知识能有效地迁移到代理任务中。 AI

影响 这种方法可以通过提高代理对其运行环境的理解能力,从而实现更强大的代理,并可能加速复杂任务自动化的进展。

排序理由 前沿实验室模型发布,附带系统卡和基准测试结果。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 8 个来源。 我们如何撰写摘要 →

阿里巴巴Qwen发布AgentWorld语言模型用于环境模拟

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Frontier Release
前沿实验室模型发布,附带系统卡和基准测试结果。
Source corroboration
8 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
106 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [8]

  1. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    🧠 Paradigm II — 智能体基础模型:世界建模作为智能体能力。

    🧠 Paradigm II — Agent Foundation Model: world modeling as agent capability. Single-turn, non-agentic environment prediction → tested directly on multi-turn, tool-calling agent tasks. No agentic RL, no task-specific tuning. Gains across 7 benchmarks, including 3 entirely https:/…

  2. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    第二部分:研究世界模型在智能体训练中的作用

    Part II: Investigating the Role of World Modeling in Agent Training 🔬 Paradigm I — Decoupled Simulation: world model as environment simulator for agent RL. The key is controllability: 1️⃣ Zero-shot generalization to 4k OOD OpenClaw environments → +4.3 Claw-Eval, +7.1 https://t…

  3. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    📊 AgentWorldBench:一个包含真实环境地面真实观测的7领域基准,由9个已建立基准上的5个前沿模型轨迹构建而成

    📊 AgentWorldBench: 7-domain benchmark with ground-truth observations from real environments, constructed from 5 frontier model trajectories on 9 established benchmarks. Results: Qwen-AgentWorld-397B-A17B achieves the highest overall score (58.71), outperforming Claude Opus 4.8 h…

  4. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    第一部分:构建环境模拟的基础模型

    Part I: Building the Foundation Model for Environment Simulation Faithful environment simulation requires multi-step causal reasoning, stateful tracking, and domain-specific knowledge. Frontier LLMs have some simulation ability from pretraining — but it's incidental, not an http…

  5. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    📣📣 认识 Qwen-AgentWorld — 一个原生语言世界模型,可在单一模式下模拟 7 种代理环境(MCP、Search、Terminal、SWE、Web、OS、Android)

    📣📣 Meet Qwen-AgentWorld — a native language world model that simulates 7 agent environments (MCP, Search, Terminal, SWE, Web, OS, Android) within a single model. Environment modeling is the training objective from day one, not a post-hoc adaptation. 🤔 LLMs are trained to be htt…

  6. arXiv cs.CL TIER_1 English(EN) · Guangfeng Cai, Kaibing Yang, Shuo He, Yu Li, Shengtian Yang, Jiaqi Lv, Lei Feng ·

    超越下一观测预测:用于序列决策的Agent自编世界模型

    arXiv:2606.25421v1 Announce Type: new Abstract: Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which…

  7. arXiv cs.CL TIER_1 English(EN) · Lei Feng ·

    超越下一观测预测:用于序列决策的Agent自编世界模型

    Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which may omit the dynamics most relevant to the agen…

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越下一观测预测:用于顺序决策的Agent自编世界模型

    Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which may omit the dynamics most relevant to the agen…