Researchers have developed a new benchmark and dataset, ECTs-100, to evaluate how well large language models (LLMs) can perform trustworthy analysis of earnings call transcripts. This benchmark focuses on groundedness, ensuring claims are supported by citations from the source documents, and correctness, assessing the accuracy of the information provided. The study found that while LLMs are proficient at grounding their analyses, they struggle with correctness and identifying when evidence is insufficient, leading to unsupported claims. AI
IMPACT This benchmark will help researchers and developers improve LLM accuracy and trustworthiness in financial analysis.
RANK_REASON This is a research paper introducing a new benchmark and dataset for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Connected Papers
- CORE Recommender
- DagsHub
- earnings call transcripts
- ECTs-100
- Gotit.pub
- Hugging Face
- Influence Flower
- large-language models
- Litmaps
- ScienceCast
- scite Smart Citations
- S&P 500
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →