A new research paper explores the reliability of large language models (LLMs) in tax reasoning systems, particularly when faced with defective inputs such as missing or contradictory facts. While LLMs can accurately compute tax liabilities on well-formed cases, their performance degrades significantly with imperfect data. The study found that models often compute through contradictions rather than abstaining, but can effectively identify these issues when prompted to verify the input. By integrating a verification step, the models' ability to abstain from defective inputs was largely recovered with minimal impact on accuracy. AI
IMPACT Highlights the need for robust error detection in LLMs beyond benchmark accuracy, crucial for reliable AI deployment in sensitive applications.
RANK_REASON Research paper published on arXiv detailing LLM capabilities and limitations in a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Large language models
- Litmaps
- ScienceCast
- scite Smart Citations
- Tax reasoning systems
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →