PulseAugur
EN
LIVE 09:55:11

New framework VibeWorlding enables AI to build 3D worlds from text

Researchers have developed VibeWorlding, a framework for training and benchmarking multimodal agents capable of constructing 3D open worlds from text prompts. The system includes a benchmark dataset (VWE-BENCH) and a reinforcement learning framework (VibeWorlding-Gym). Experiments indicate that current large language models, including GPT-5.5 and Qwen3.8-Max, struggle with this task, achieving less than 60% success. However, reinforcement learning training significantly improves open-source models, with VibeWorlder-30B-A3B outperforming even frontier models in performance. AI

IMPACT This research advances agentic capabilities in 3D content creation, potentially impacting fields like game development and virtual reality.

RANK_REASON The cluster describes a research paper introducing a new framework and benchmark for multimodal agents in 3D world generation.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

New framework VibeWorlding enables AI to build 3D worlds from text

COVERAGE [4]

  1. arXiv cs.AI TIER_1 English(EN) · Yansong Ning, Jingwen Ye, Zhongkai Wu, Yang Sun, Yiqin Zhu, Xingyi Li, Weidong Zhang, Hao Liu ·

    VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

    arXiv:2608.15265v1 Announce Type: new Abstract: Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agent…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

    A unified framework benchmarks and trains multimodal agents that infer intent, plan 3D scenes, invoke tools, and reflect on feedback, revealing that reinforcement learning improves open-source models beyond closed-source frontiers.

  3. arXiv cs.CV TIER_1 English(EN) · Shengyuan Wang, Zhiheng Zheng, Yu Shang, Lixuan He, Yangcheng Yu, Fan Hangyu, Jie Feng, Qingmin Liao, Yong Li ·

    UrbanWorld2.0: A Multimodal Agentic Framework for Reality-Aligned 3D World Generation at City-Scale

    arXiv:2511.18005v2 Announce Type: replace Abstract: The automated generation of high-fidelity, city-scale 3D environments remains a formidable challenge with profound academic and industrial implications. However, existing methods struggle to achieve the necessary quality, fideli…

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Multimodal agents are building entire 3D worlds from scratch. VibeWorlding turned text prompts into open worlds—end-to-end, with 20 upvotes on Hugging Face. The

    Multimodal agents are building entire 3D worlds from scratch. VibeWorlding turned text prompts into open worlds—end-to-end, with 20 upvotes on Hugging Face. The paper: https:// huggingface.co/papers/2608.152 65 # AI # MachineLearning # Research