PulseAugur
EN
LIVE 08:50:44
ENTITY WorldBench

WorldBench

PulseAugur coverage of WorldBench — every cluster mentioning WorldBench across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
3
3 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
3
3 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

1 day(s) with sentiment data

RECENT · PAGE 1/1 · 3 TOTAL
  1. TOOL · CL_289311 ·

    New WorldBench benchmark evaluates LLMs on 3D world generation

    Researchers have introduced WorldBench, a new benchmark designed to evaluate large language models' ability to generate interactive 3D worlds using Three.js. Current evaluation methods, relying on either visual snapshot…

  2. RESEARCH · CL_221112 ·

    New benchmarks reveal LLM agent limitations in multilingual tasks and collaboration

    New research explores the capabilities of large language model (LLM) agents in complex, multilingual, and collaborative environments. WorldBench, a new benchmark, tests LLM agents across 1,600 tasks in seven languages a…

  3. TOOL · CL_208692 ·

    New WorldBench benchmark isolates physics concepts for AI world model evaluation

    Researchers have introduced WorldBench, a new benchmark designed to evaluate the physical understanding of world models used in AI. Unlike previous benchmarks that test multiple physics concepts simultaneously, WorldBen…