McNemar
PulseAugur coverage of McNemar — every cluster mentioning McNemar across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
AI agents in medical imaging show higher deference to human-attributed false findings
A pilot audit examined how AI agents used in medical imaging respond to falsified findings, specifically whether they retract correct answers when presented with incorrect information. The study found that agents were s…
-
New AI method trains medical triage agents using existing guidelines
Researchers have developed a novel method called Guideline-as-Oracle (GAO) to train ophthalmic telephone triage agents with minimal human annotation. GAO compiles existing clinical guidance from the American Academy of …
-
AI benchmark audit reveals reproducibility issues, prompts withdrawal of claims
A forensic audit of a radiology vision-language model benchmark revealed significant discrepancies between its intended protocol and the released artifacts. The audit found issues with DICOM rendering, dataset splitting…
-
AI equity forecasting benchmark reveals LoRA-adapted TimesFM lacks directional skill
A new research paper challenges the effectiveness of large language models like TimesFM for equity forecasting, particularly when using LoRA adapters. The study introduces a base-rate-honest benchmark to expose how seem…
-
New Benchmark Reveals LLM Reasoning Decay Under Depth Scaling
Researchers have introduced the Complexity Ceiling Benchmark (CCB) to evaluate how language models' sequential reasoning abilities degrade with increasing task depth. Across six thousand trials involving five frontier a…