PulseAugur
实时 06:28:36
English(EN) OpenSanctions Pairs: Large-Scale Entity Matching with LLMs

发布新的基准数据集 OpenSanctions Pairs,GPT-4o 在实体匹配性能上领先

发布了一个名为 OpenSanctions Pairs 的新基准数据集,专为制裁和 OSINT 数据的大规模实体匹配而设计。该数据集包含超过 755,000 个专家标记的配对,源自 293 个来源和 45 个司法管辖区的 100 万多个实体,具有显著的语言和结构多样性。评估显示,GPT-4o 取得了最高性能,F1 分数为 99.0%,紧随其后的是开源模型 DeepSeek-R1-Distill-Qwen-14B,F1 分数为 98.2%。这些结果表明,当前的匹配性能已接近其实际极限,重点将转向阻塞和聚类等其他管道组件。 AI

影响 为实体匹配基准设定了新标准,将大型语言模型在合规和 OSINT 数据处理方面的性能上限推向了新的高度。

排序理由 该集群是关于一篇介绍基准数据集并评估大型语言模型在实体匹配任务上性能的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

发布新的基准数据集 OpenSanctions Pairs,GPT-4o 在实体匹配性能上领先

本文如何被排名

Signal score
30 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群是关于一篇介绍基准数据集并评估大型语言模型在实体匹配任务上性能的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Chandler Smith, Magnus Sesodia, Friedrich Lindenberg, Christian Schroeder de Witt ·

    OpenSanctions 配对:使用 LLM 进行大规模实体匹配

    arXiv:2603.11051v2 Announce Type: replace-cross Abstract: We release OpenSanctions Pairs, the first large-scale public benchmark for entity matching on sanctions and OSINT data. The dataset includes 755,540 expert-labeled pairs over 1 million entities, aggregated from 293 source …