Researchers have developed a method to stress-test Large Language Models (LLMs) like GPT-4.1 mini in their ability to correctly cite financial evidence. The study focuses on a problem where LLMs can produce numerically correct calculations but cite the wrong financial roles or sources. By swapping citations between cells with identical numbers, the researchers isolated the LLM's role-recognition capabilities, revealing a trade-off between detecting incorrect citations and supporting valid ones. This evaluation framework aims to make numerical correctness, cited-role support, and acceptance outcomes independently assessable for LLM-based financial assistants. AI
IMPACT This research could lead to more reliable LLM-based financial analysis tools by improving their ability to accurately cite evidence.
RANK_REASON The cluster contains a research paper detailing a new evaluation method for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- GPT-4.1 mini
- Hugging Face
- Jevíčko
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →