PulseAugur
EN
LIVE 13:30:07
ENTITY Cohen's kappa

Cohen's kappa

PulseAugur coverage of Cohen's kappa — every cluster mentioning Cohen's kappa across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
7
15 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
7
14 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

6 day(s) with sentiment data

RECENT · PAGE 1/1 · 15 TOTAL
  1. TOOL · CL_193738 ·

    LLM A/B test prediction struggles with reliability, study finds

    A new research paper explores the effectiveness of large language models in predicting the outcomes of A/B tests for web page designs. The study found that while a Gemini 3 Flash model could achieve a moderate agreement…

  2. TOOL · CL_183253 ·

    New VIVID benchmark reveals AI's figurative language gap in Vietnamese

    Researchers have introduced VIVID, a new benchmark designed to assess how well AI models understand figurative language within the Vietnamese language and culture. The benchmark includes over 1,600 idioms and proverbs, …

  3. RESEARCH · CL_184907 ·

    LLMs fabricate user profiles, new research finds · 2 sources tracked

    A new research paper introduces MirageBench, a dataset and evaluation framework to study how large language models (LLMs) fabricate user attributes, a phenomenon termed over-inference (OI). The study found that all 12 e…

  4. TOOL · CL_179892 ·

    New paper categorizes AI agent failure modes and enables automated tracking

    A new paper introduces a framework for categorizing 41 agent failure modes based on their origin within the interaction between components like models, harnesses, and users. This approach attributes bugs to the 'seams' …

  5. TOOL · CL_160820 ·

    LLM framework enhances accuracy in identifying adverse drug events

    A new research paper details a human-in-the-loop framework utilizing a retrieval-augmented, multi-agent large language model (LLM) to identify cutaneous immune-related adverse events (cirAEs) from clinical notes. This L…

  6. TOOL · CL_147900 ·

    New Analytic Abduction Framework Enhances Human-AI Coordination

    Researchers have introduced Analytic Abduction, a novel framework for human-AI coordination that focuses on the analytic mode of abductive reasoning. This approach identifies latent factors contributing to complex obser…

  7. RESEARCH · CL_143397 ·

    New framework uses VLMs to improve EEG-to-image reconstruction evaluation

    Researchers have developed a new framework to evaluate the coherence between EEG signals and reconstructed images, addressing limitations in existing metrics like SSIM and LPIPS. This framework utilizes four Vision-Lang…

  8. RESEARCH · CL_133121 ·

    SynthAVE uses LLM arena for scalable e-commerce data labeling · 2 sources tracked

    Researchers have developed SynthAVE, a novel system for generating and validating synthetic labels for e-commerce attribute extraction at an industrial scale. This approach addresses the prohibitive cost of human labeli…

  9. RESEARCH · CL_106950 ·

    LLM-as-judge tools fail to prioritize human validation, study finds

    A recent evaluation of six LLM-as-judge tools revealed that most prioritize generating scores over ensuring the trustworthiness of those scores. The author argues that a judge's validation against human labels, measured…

  10. RESEARCH · CL_105153 ·

    LLMs analyzed for self-stigma support in drug use communities · 2 sources tracked

    Researchers have developed methods to analyze self-stigma expressed by individuals who use drugs in online communities, specifically on Reddit. One study created a codebook to categorize self-stigma into cognitive, affe…

  11. TOOL · CL_100060 ·

    New framework measures university CS curriculum alignment with global standards

    A new framework has been developed to measure how well university computer science programs align with international curricular guidelines, specifically CS2013 and CS2023. This human-in-the-loop pipeline represents prog…

  12. RESEARCH · CL_99671 ·

    LLM-as-a-Judge models show significant reliability and bias issues, study finds

    A new study evaluating LLM-as-a-Judge models reveals significant issues with their reliability and validity. The research, which analyzed 21 judges across multiple benchmarks and over 541,000 judgments, found that commo…

  13. TOOL · CL_93144 ·

    LLMs show promise in identifying discourse units for aphasia assessment

    A new research paper explores the use of instruction-tuned large language models (LLMs) for classifying Correct Information Units (CIUs) in aphasic discourse. The study found that while zero-shot prompting was insuffici…

  14. TOOL · CL_52901 ·

    LLM judge evaluations require hundreds of labels for reliable results

    A recent article highlights the critical need for larger evaluation datasets when using LLMs as judges in AI model assessments. The author explains that common practice of using small, ad-hoc datasets is insufficient fo…

  15. TOOL · CL_18536 ·

    LLM system aids explainable defect analysis in laser powder bed fusion

    Researchers have developed a new decision-support system that combines structured knowledge about defects with large language models (LLMs) to analyze and guide mitigation strategies in laser powder bed fusion (LPBF) ma…