A new study published on arXiv reveals that chemical reasoning language models, despite being expected to provide accurate molecular answers through chain-of-thought (CoT) reasoning, frequently exhibit widespread hallucination. This fabrication of structural claims is often decoupled from the correctness of the final answer. The research indicates that CoT serves as a "hallucination-prone molecular scratchpad," where models like Chem-R and ether-0 utilize fragmented SMILES drafts, and ChemDFM-R emphasizes scaffold and naming cues. Perturbing these structural drafts in Chem-R demonstrated a causal impact on generation, suggesting that these internal molecular representations are load-bearing even when verbal claims are not. AI
IMPACT Highlights the unreliability of chain-of-thought reasoning in chemical models, cautioning against direct interpretation of CoT as faithful reasoning and motivating process-level supervision.
RANK_REASON Academic paper detailing findings about language model reasoning processes. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →