QuanTemp is a new benchmark designed to evaluate fact-checking models on their ability to handle numerical claims. It comprises 15,514 numerical claims sourced from 45 fact-checkers and a corpus of 423,320 snippets. The benchmark aims to distinguish between genuine improvements in numerical reasoning and mere luck from better search results, though current top models achieve a 58.32 macro-F1 score. AI
IMPACT This benchmark could lead to more robust AI fact-checking systems capable of handling numerical data accurately.
RANK_REASON The cluster describes a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →