Researchers have introduced DelistBench, a new benchmark designed to evaluate search-enabled large language models (LLMs) for their ability to accurately complete corporate-event databases. The benchmark, comprising 1,200 security-level delisting announcements, aims to help financial institutions independently verify missing, stale, or misclassified records. Evaluations showed that web access significantly improves LLMs' accuracy in identifying announcement dates and event statuses, with cost-effective systems achieving competitive performance. AI
IMPACT This benchmark could improve the accuracy and efficiency of financial data verification by LLMs.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- DelistBench
- Gotit.pub
- Hugging Face
- LLMs
- ScienceCast
- Search-to-Record
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →