A new audit of financial news NLP benchmarks reveals significant temporal leakage, where random train-test splits inflate performance metrics by up to 6.5x compared to chronological splits. This leakage is particularly pronounced with larger models and richer features. The study found that only mergers and acquisitions (M&A) news showed a positive signal under near-temporal chronological evaluation, but this signal was localized to specific semantic contexts and did not transfer to broader datasets. Researchers advocate for leakage audits as a mandatory disclosure for financial NLP benchmarks. AI
IMPACT Highlights critical flaws in financial NLP benchmark evaluation, necessitating stricter auditing practices for reliable model performance assessment.
RANK_REASON The item is an academic paper detailing a new audit methodology and findings for NLP benchmarks. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- DeBERTa-v3-large
- FinBERT
- FNSPID
- Llama 3
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- Qwen2.5
- RoBERTa-large
- tf–idf
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →