A new benchmark dataset called OpenSanctions Pairs has been released, designed for large-scale entity matching specifically for sanctions and OSINT data. The dataset contains over 755,000 expert-labeled pairs derived from more than 1 million entities across 293 sources and 45 jurisdictions, offering significant linguistic and structural diversity. Evaluations show that GPT-4o achieved the highest performance with a 99.0% F1 score, closely followed by an open-source model, DeepSeek-R1-Distill-Qwen-14B, at 98.2% F1. These results suggest that current matching performance is nearing its practical limits, shifting focus to other pipeline components like blocking and clustering. AI
IMPACT Sets a new standard for entity matching benchmarks, pushing the performance ceiling for LLMs in compliance and OSINT data processing.
RANK_REASON The cluster is about a new academic paper introducing a benchmark dataset and evaluating LLMs on entity matching tasks. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- DeepSeek-R1-Distill-Qwen-14B
- GPT-4o
- Magnus Sesodia
- MIPROv2
- nomenklatura RegressionV1
- OpenSanctions Pairs
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →