PulseAugur
实时 04:14:49
English(EN) CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

新的CorporateBench基准测试揭示LLM在处理大规模企业数据时遇到困难

一项名为CorporateBench (CB) 的新基准测试已被引入,用于评估大型语言模型 (LLM) 在回答来自大型、随时间演变的企业文档集合中的复杂问题的能力。CB旨在解决现有基准测试的局限性,包含超过230,000份文档,并在信息提取和知识库查询方面评估LLM。对五种LLM的初步评估显示,当输入规模接近实际企业规模时,性能显著下降,这凸显了当前LLM在企业沟通方面推理能力的重大差距。 AI

影响 凸显了LLM在处理企业数据方面的性能差距,可能推动开发更适合企业知识管理的模型。

排序理由 该集群描述了一篇介绍用于评估LLM的新型基准测试的学术论文。

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的CorporateBench基准测试揭示LLM在处理大规模企业数据时遇到困难

本文如何被排名

Signal score
6 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一篇介绍用于评估LLM的新型基准测试的学术论文。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Igor Labutov ·

    CorporateBench:具有时间知识库的大规模问答基准测试

    LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated mul…