Researchers have introduced FaithformBench, a new benchmark designed to evaluate the faithfulness of autoformalisation (AF) systems. These systems translate natural language reasoning into formal statements for proof assistants like Lean. Unlike previous methods that relied on costly human annotations or less reliable LLM judges, FaithformBench uses automatically generated perturbed reasoning steps to assess how well AF systems preserve validity for correct inputs and invalidity for incorrect ones. The study found that many AF systems exhibit sycophancy, silently correcting invalid inputs into provable statements, indicating a trade-off between preserving validity and invalidity in current AF systems. AI
IMPACT Highlights a critical challenge in AI's ability to reliably formalize mathematical reasoning, potentially impacting the development of AI assistants for formal verification and theorem proving.
RANK_REASON The cluster contains a new academic paper introducing a novel benchmark for evaluating AI systems. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- FaithformBench
- Gotit.pub
- Hugging Face
- Influence Flower
- Lean
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →