PulseAugur
实时 06:45:16
English(EN) Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency

LLM筛选工作流在研究综合中表现不稳定 · arXiv论文

一篇新发表在arXiv上的研究评估了大型语言模型(LLM)在筛选研究论文以进行证据综合方面的有效性。研究发现,虽然没有工作流(无论是人类还是LLM)能够识别所有相关研究,但LLM的表现因处理配置而非模型本身而存在显著差异。研究表明,由于运行间一致性等问题,LLM最适合用于经过验证的、人类监督下的高召回率场景工作流,而不是自主排除。 AI

影响 LLM筛选性能高度依赖于工作流配置,这表明需要谨慎地将其整合到人类监督的研究过程中。

排序理由 该集群包含一篇详细介绍LLM性能研究结果的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LLM筛选工作流在研究综合中表现不稳定 · arXiv论文

本文如何被排名

Signal score
27 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍LLM性能研究结果的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Nikol Figalov\'a, Lynn Huestegge, Anne B\"ockler-Raettig ·

    概念复杂范围审查中人类和LLM筛选工作流程的评估:召回率-工作量权衡与运行间一致性

    arXiv:2608.26885v1 Announce Type: new Abstract: Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screenin…