A new benchmark called KINA has been introduced to evaluate large language models across 261 fine-grained disciplines, addressing issues of scaling-driven design and annotation quality. The benchmark, comprising 899 items, was used to evaluate 42 models from 13 different labs. Gemini-3.1-Pro-Preview emerged as the top performer with a score of 53.17%, followed by Claude-Opus-4.6 and GPT-5.4, indicating substantial room for improvement across models. AI
IMPACT Establishes a new evaluation standard for LLMs, highlighting performance tiers and the impact of tool augmentation.
RANK_REASON The cluster contains a research paper introducing a new benchmark for LLMs and reporting evaluation results.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →