PulseAugur
EN
LIVE 16:05:42
ENTITY Cohen's kappa

Cohen's kappa

PulseAugur coverage of Cohen's kappa — every cluster mentioning Cohen's kappa across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
8
19 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
6
17 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

6 day(s) with sentiment data

RECENT · PAGE 1/2 · 26 TOTAL
  1. TOOL · CL_254607 ·

    LLM-derived weights enhance psychometric model for educational assessment reliability

    Researchers have developed a new Rater Ising-Potts model that leverages Large Language Model (LLM) embeddings to assess the reliability of educational assessments. This model focuses on pairwise agreement between rating…

  2. TOOL · CL_252064 ·

    New protocol evaluates AI tutors' response to disengaged students

    Researchers have developed a new protocol called Disengagement-Aware Student Simulators (DAS2) to evaluate AI tutors by modeling five learner-engagement states: engaged, gaming, wheel-spinning, off-task, and mixed. This…

  3. COMMENTARY · CL_249144 ·

    Human evaluation remains critical for LLM quality assessment

    Human evaluation is crucial for assessing Large Language Models (LLMs) because automated metrics like BLEU scores often fail to capture nuanced qualities such as coherence, creativity, and factual accuracy. This approac…

  4. TOOL · CL_247864 ·

    New Supervised Hebbian Learning Algorithm for Spiking Neural Networks Outperforms STDP

    Researchers have developed a new gradient-free supervised learning algorithm for spiking neural networks (SNNs) called Supervised Spike Agreement-Dependent Plasticity (Supervised SADP). This method directly embeds class…

  5. COMMENTARY · CL_244659 ·

    Multi-agent AI failures stem from process, not models, study finds

    Multi-agent AI systems frequently fail not due to model limitations, but due to poorly defined processes and inter-agent communication issues. A study analyzing over 1,600 execution traces found that approximately 80% o…

  6. TOOL · CL_229166 ·

    Study: LLMs show limited agreement with human persuasion judgments

    A new study published on arXiv investigates whether large language models (LLMs) update their beliefs in response to persuasive arguments in a manner similar to humans. The research found that LLMs exhibit only slight a…

  7. TOOL · CL_228616 ·

    Frontier LLMs fail oncology decision-making benchmark, new study finds

    A new benchmark, the Oncology Decision Boundary Benchmark (ODBB), has been developed to evaluate the decision-making capabilities of frontier large language models (LLMs) in oncology. The study found that even advanced …

  8. RESEARCH · CL_227148 ·

    AI models in pathology gain robustness through new benchmarking and artifact generation techniques

    Two new research papers explore methods for improving the reliability and robustness of AI models in computational pathology. The first paper, "Reliable Benchmarking of Artifact Detection in Computational Pathology," pr…

  9. TOOL · CL_217822 ·

    New metrics proposed for imbalanced classification problems

    A new research paper introduces robust modifications to common performance metrics used in imbalanced classification problems. The authors demonstrate that existing metrics like Matthews' correlation coefficient (MCC) a…

  10. TOOL · CL_205795 ·

    New ARISE method enhances feature selection for small-sample biomedical omics data

    Researchers have developed ARISE, a novel ensemble method for feature selection in small-sample biomedical omics data. This adaptive framework integrates multiple relevance signals, assesses class-balanced stability, an…

  11. TOOL · CL_193738 ·

    LLM A/B test prediction struggles with reliability, study finds

    A new research paper explores the effectiveness of large language models in predicting the outcomes of A/B tests for web page designs. The study found that while a Gemini 3 Flash model could achieve a moderate agreement…

  12. TOOL · CL_201679 ·

    Bangla math reasoning in LLMs: CoT supervision shows mixed results

    A new study, MathShikkha, investigated the effectiveness of Chain-of-Thought (CoT) supervision for improving mathematical reasoning in small language models trained on the Bangla language. The research constructed a Ban…

  13. TOOL · CL_183253 ·

    New VIVID benchmark reveals AI's figurative language gap in Vietnamese

    Researchers have introduced VIVID, a new benchmark designed to assess how well AI models understand figurative language within the Vietnamese language and culture. The benchmark includes over 1,600 idioms and proverbs, …

  14. RESEARCH · CL_184907 ·

    LLMs fabricate user profiles, new research finds · 2 sources tracked

    A new research paper introduces MirageBench, a dataset and evaluation framework to study how large language models (LLMs) fabricate user attributes, a phenomenon termed over-inference (OI). The study found that all 12 e…

  15. TOOL · CL_179892 ·

    New paper categorizes AI agent failure modes and enables automated tracking

    A new paper introduces a framework for categorizing 41 agent failure modes based on their origin within the interaction between components like models, harnesses, and users. This approach attributes bugs to the 'seams' …

  16. TOOL · CL_160820 ·

    LLM framework enhances accuracy in identifying adverse drug events

    A new research paper details a human-in-the-loop framework utilizing a retrieval-augmented, multi-agent large language model (LLM) to identify cutaneous immune-related adverse events (cirAEs) from clinical notes. This L…

  17. TOOL · CL_147900 ·

    New Analytic Abduction Framework Enhances Human-AI Coordination

    Researchers have introduced Analytic Abduction, a novel framework for human-AI coordination that focuses on the analytic mode of abductive reasoning. This approach identifies latent factors contributing to complex obser…

  18. RESEARCH · CL_143397 ·

    New framework uses VLMs to improve EEG-to-image reconstruction evaluation

    Researchers have developed a new framework to evaluate the coherence between EEG signals and reconstructed images, addressing limitations in existing metrics like SSIM and LPIPS. This framework utilizes four Vision-Lang…

  19. RESEARCH · CL_133121 ·

    SynthAVE uses LLM arena for scalable e-commerce data labeling · 2 sources tracked

    Researchers have developed SynthAVE, a novel system for generating and validating synthetic labels for e-commerce attribute extraction at an industrial scale. This approach addresses the prohibitive cost of human labeli…

  20. RESEARCH · CL_106950 ·

    LLM-as-judge tools fail to prioritize human validation, study finds

    A recent evaluation of six LLM-as-judge tools revealed that most prioritize generating scores over ensuring the trustworthiness of those scores. The author argues that a judge's validation against human labels, measured…