PulseAugur
EN
LIVE 09:28:49

Chemical CoT models hallucinate molecular structures, acting as prone scratchpads

A new study published on arXiv reveals that chemical reasoning language models, despite being expected to provide accurate molecular answers through chain-of-thought (CoT) reasoning, frequently exhibit widespread hallucination. This fabrication of structural claims is often decoupled from the correctness of the final answer. The research indicates that CoT serves as a "hallucination-prone molecular scratchpad," where models like Chem-R and ether-0 utilize fragmented SMILES drafts, and ChemDFM-R emphasizes scaffold and naming cues. Perturbing these structural drafts in Chem-R demonstrated a causal impact on generation, suggesting that these internal molecular representations are load-bearing even when verbal claims are not. AI

IMPACT Highlights the unreliability of chain-of-thought reasoning in chemical models, cautioning against direct interpretation of CoT as faithful reasoning and motivating process-level supervision.

RANK_REASON Academic paper detailing findings about language model reasoning processes. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Chemical CoT models hallucinate molecular structures, acting as prone scratchpads

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Jiatong Li, Yuxuan Ren, Weida Wang, Xiaoyong Wei, Yatao Bian ·

    Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

    arXiv:2607.20935v1 Announce Type: cross Abstract: Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning model families and twelve chemistry tasks, hallucination is widespread and largely…