Researchers have introduced TeXFix-Bench, a new benchmark designed to evaluate Large Language Models (LLMs) in their ability to repair errors in document source code. This benchmark is grounded in an empirical taxonomy of faults derived from real-world sources like TeX Stack Exchange and GitHub commits, covering formats such as LaTeX, Typst, and Markdown. Evaluations using TeXFix-Bench revealed significant differences in LLM performance, with Typst proving more challenging to repair than LaTeX and Markdown, and highlighting that successful compilation does not always equate to high-quality content restoration. AI
IMPACT This benchmark could lead to more robust LLMs for technical writing and code generation tasks.
RANK_REASON The cluster contains a research paper introducing a new benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →