Human annotators
PulseAugur coverage of Human annotators — every cluster mentioning Human annotators across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Task decomposition ineffective for LLM-based NLG evaluation, study finds
A new research paper challenges the effectiveness of task decomposition in improving Natural Language Generation (NLG) evaluation using the LLM-as-a-Judge framework. The study found no performance gains from decompositi…
-
New framework analyzes LLM bias in content moderation
Researchers have developed a new framework called the Ghost Annotator to analyze human label variation in content moderation tasks, particularly when LLMs are used for annotation. This framework combines conformal predi…
-
New benchmark reveals AI visual misinformation detection failures
Researchers have developed SynCred-Bench, a new benchmark designed to evaluate the detection of AI-generated visual misinformation that mimics credible sources. The benchmark includes 600 AI-generated images and a set o…
-
LLMs improve NLI dataset error detection and model fine-tuning
A new framework called EVADE uses large language models (LLMs) to generate and validate explanations for error detection in natural language inference (NLI) datasets. This approach aims to reduce the cost and effort ass…