Researchers have developed ORQA, a novel framework designed to evaluate the occupation-specific knowledge of large language models. This method connects O*NET occupations to authoritative websites, generating source-traceable question-answer pairs. The framework covers 116 occupations and includes 480 questions derived from 187 websites, with a focus on real-world skill relevance. In testing, Claude Opus 4.6, GPT-5.4, and Claude Sonnet 4.6 demonstrated the highest performance, achieving around 58-62% accuracy, while smaller open-weight models performed at 33-41%. AI
IMPACT Establishes a new benchmark for evaluating LLM professional knowledge, highlighting performance disparities across occupations and models.
RANK_REASON The cluster describes a new research paper introducing a framework for evaluating LLM knowledge in professional domains. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude Opus 4.6
- Claude Sonnet 4.6
- Fish and Game Wardens
- GPT-5.4
- O*NET OnLine
- Sheet Metal Workers
- Shreyas Krishnan
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →