Researchers have developed DiscoverPhysics, a new benchmark designed to test the scientific reasoning capabilities of large language models (LLMs). This interactive benchmark challenges LLM agents to discover the laws of motion in simulated worlds with novel physics, requiring them to design experiments, observe data, and propose natural-language explanations and Python implementations of inferred laws. Evaluations across eleven frontier models revealed that even the strongest agents pass only half the worlds, particularly struggling with uncovering latent structures. The benchmark also highlighted a significant gap between open-source and commercial models, with the latter demonstrating superior experimental design and data extraction abilities. AI
IMPACT This benchmark could drive the development of LLMs with more robust scientific reasoning and hypothesis-generation abilities.
RANK_REASON The cluster describes a new academic paper introducing a novel benchmark for evaluating LLM capabilities.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →