PulseAugur
EN
LIVE 20:57:38

LLMs struggle to reproduce physics experiment results, failing numerical simulations

A new preprint from Peking University evaluated the ability of large language models to reproduce numerical results from experimental physics papers. Researchers found that all tested LLMs, including OpenAI Codex powered by GPT-5.3, achieved a 0% end-to-end callback rate, meaning they could not replicate any full numerical outcomes. While the models demonstrated strong comprehension of the papers' methodologies, they consistently made errors in data analysis and numerical simulation, leading to incorrect final results. The study identified several failure modes, such as formula implementation errors and oversimplification of complex physical models. AI

IMPACT LLMs struggle with complex numerical simulation and data analysis in scientific research, indicating limitations beyond text comprehension.

RANK_REASON Academic paper evaluating LLM capabilities on a new domain (physics simulation).

Read on LessWrong (AI tag) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LLMs struggle to reproduce physics experiment results, failing numerical simulations

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Academic paper evaluating LLM capabilities on a new domain (physics simulation).
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
152 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. LessWrong (AI tag) TIER_1 English(EN) · fessus ·

    AI Is Bad at Physics

    <p><span>There’s a </span><a href="https://arxiv.org/pdf/2603.27646"><span>new preprint</span></a><span> from Peking University in China that assesses LLM capabilities in reproducing results from experimental physics papers. Their finding? All the agents had a </span><b><span>0% …