PulseAugur
实时 07:15:59
English(EN) ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge

新的ORQA框架测试LLM职业知识;Claude和GPT领先

研究人员开发了ORQA,一个旨在评估大型语言模型职业特定知识的新框架。该方法将O*NET职业与权威网站连接起来,生成可追溯来源的问答对。该框架涵盖116种职业,包含来自187个网站的480个问题,重点关注现实世界的技能相关性。在测试中,Claude Opus 4.6、GPT-5.4和Claude Sonnet 4.6表现最佳,准确率约为58-62%,而较小的开源模型表现为33-41%。 AI

影响 为评估LLM专业知识建立了新的基准,突显了不同职业和模型之间的性能差异。

排序理由 该集群描述了一篇介绍用于评估LLM在专业领域知识的框架的新研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的ORQA框架测试LLM职业知识;Claude和GPT领先

本文如何被排名

Signal score
23 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一篇介绍用于评估LLM在专业领域知识的框架的新研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Shreyas Krishnan, Serina Chang, Abhishek Nagaraj ·

    ORQA:一个面向LLM专业知识的职业现实问答框架

    arXiv:2609.12366v1 Announce Type: new Abstract: We present ORQA, a method for testing occupation-level knowledge in large language models. Prior methods either map abstract LLM skills to occupations via task definitions or utilize expert knowledge which is difficult to obtain at …