AVeriTeC
PulseAugur coverage of AVeriTeC — every cluster mentioning AVeriTeC across labs, papers, and developer communities, ranked by signal.
-
Automated fact-checkers struggle with numerical claims in AVeriTeC task
The second AVeriTeC shared task evaluated seven open-weight fact-checking systems. These systems operated with a 23 GB GPU and processed one claim per minute against a fixed evidence corpus. Numerical claims proved to b…
-
New method detects misinformation by analyzing LLM internal representations
Researchers have developed a novel method for detecting misinformation by analyzing the internal representations of language models, rather than relying on external knowledge or surface-level text features. This approac…
-
New AI workflow enhances claim-evidence traceability in writing
Researchers have developed a new workflow called evidence-ledger adjudication to improve the traceability of claims made by AI agents in relation to supporting evidence. This system pairs each claim with an evidence pac…
-
Fact-checking benchmarks still contaminated despite dynamic evaluation
A new study examines contamination risks in dynamic evaluation benchmarks for multimodal automated fact-checking (MAFC). The research found that even benchmarks designed with post-knowledge cut-off claims still suffer f…
-
Fact-checking benchmarks flawed by 'contamination', study finds · 2 sources tracked
A new research paper challenges the effectiveness of current benchmarks for evaluating multimodal automated fact-checking (MAFC) systems. The study reveals that even dynamic benchmarks, which use claims published after …