McNemar tests
PulseAugur coverage of McNemar tests — every cluster mentioning McNemar tests across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New ExecRetrieval benchmark reveals functional-correctness gap in code retrieval
A new benchmark called ExecRetrieval has been introduced to measure the functional correctness of code retrieval systems. The benchmark consists of 939 Python tasks, each with a correct implementation and up to four exe…
-
New RAG evaluation methods emerge for Turkish and domain-specific data · 4 sources tracked
Researchers are developing new methods to evaluate and improve Retrieval-Augmented Generation (RAG) systems. One study compares different chunking and embedding strategies for Turkish RAG, finding that layout-aware chun…
-
LLMs struggle with fine-grained emotion recognition in zero-shot tests
A new research paper evaluates the zero-shot emotion recognition capabilities of three leading large language models: Claude Sonnet 4.6, ChatGPT (GPT-5.4), and Gemini 2.5-Flash. The study found that Gemini achieved the …
-
New benchmark reveals Vision-Language Models struggle with script consistency
A new benchmark, PuMVR, has been developed to evaluate Vision-Language Models (VLMs) on their ability to handle multiple scripts within a single language. The benchmark, comprising 1,000 parallel image-text instances ac…