PulseAugur
EN
LIVE 14:20:43
ENTITY Rouge

Rouge

PulseAugur coverage of Rouge — every cluster mentioning Rouge across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
7
16 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
6
15 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

5 day(s) with sentiment data

RECENT · PAGE 1/1 · 16 TOTAL
  1. TOOL · CL_195467 ·

    BLEU and ROUGE metrics explained for language model evaluation

    BLEU and ROUGE are key metrics used to evaluate the performance of language models, particularly in tasks like machine translation and text summarization. BLEU focuses on precision of n-grams and includes a penalty for …

  2. TOOL · CL_183936 ·

    LLM-as-a-Judge: Using AI to Evaluate AI Output

    The "LLM-as-a-Judge" technique utilizes a large language model to evaluate the output of other models, addressing the bottleneck of performance assessment in AI development. This method acts as a scalable and explainabl…

  3. RESEARCH · CL_184949 ·

    MIDAS framework enhances enterprise text summarization with multi-LLM adaptation

    Researchers have introduced MIDAS (Multi-LLM Iterative Data-Adaptive Summarization), a novel framework designed to enhance text summarization for enterprise applications. MIDAS utilizes a multi-LLM approach that incorpo…

  4. TOOL · CL_174027 ·

    AI-generated counterspeech can be personalized for greater impact

    Researchers have developed and evaluated strategies for generating contextualized AI-powered counterspeech to combat online toxicity. Unlike generic approaches, these new methods adapt to the conversational context and …

  5. TOOL · CL_156439 ·

    New benchmark PathReportEval standardizes pathology report generation evaluation

    Researchers have introduced PathReportEval, a new benchmark and evaluation framework designed to standardize the assessment of pathology report generation from whole-slide images. This framework addresses the limitation…

  6. TOOL · CL_145676 ·

    Small language models show strong biomedical text generation after alignment

    A new research paper explores post-training alignment techniques for small language models (SLMs) specifically for biomedical data-to-text generation. The study compares supervised fine-tuning (SFT), Direct Preference O…

  7. TOOL · CL_143804 ·

    New research paper calls for improved evaluation of personalized dialogue systems

    A new research paper published on arXiv proposes a shift in how retrieval-augmented personalized dialogue systems are evaluated. The study highlights that current metrics like BLEU, ROUGE, and F1 fail to capture the dee…

  8. TOOL · CL_117473 ·

    Customized Generative AI Agents Developed for Transportation Engineering

    Researchers have developed a method for customizing generative AI agents for specialized fields like transportation engineering. By using a curated dataset of U.S. transportation documents, they fine-tuned six large lan…

  9. TOOL · CL_115619 ·

    AI researchers call for stricter terminology in machine unlearning for LLMs

    A position paper argues that the term "machine unlearning" is frequently misused in the context of large language models (LLMs). The authors propose that "machine unlearning" should strictly refer to the process of remo…

  10. RESEARCH · CL_109560 ·

    New defense framework targets data poisoning in text summarization models

    Researchers have developed a new framework to defend text summarization models against data poisoning attacks that occur during the fine-tuning stage. This method, called Detect, Unlearn, Restore, can identify poisoned …

  11. RESEARCH · CL_109567 ·

    Fine-tuned PEGASUS model achieves state-of-the-art abstractive summarization

    Researchers have fine-tuned the PEGASUS model on the XL-Sum English corpus to improve abstractive summarization performance. This fine-tuned model achieved state-of-the-art results on the XL-Sum English Corpus, demonstr…

  12. RESEARCH · CL_86679 ·

    Direct Preference Optimization Simplifies LLM Fine-Tuning

    Researchers have published a study on Direct Preference Optimization (DPO), a reinforcement learning technique for fine-tuning large language models. The paper details how DPO simplifies training, enhances computational…

  13. TOOL · CL_78028 ·

    LLM-as-a-Judge replaces traditional metrics for AI evaluation

    Traditional NLP metrics like BLEU and ROUGE are insufficient for evaluating generative AI responses in production, especially in complex domains like financial regulatory documentation. These metrics, designed for tasks…

  14. RESEARCH · CL_14134 ·

    New RCD method optimizes LLM processing of long clinical texts within budget

    Researchers have developed a new method called RCD for selecting relevant subsets of long clinical texts to reduce token costs for large language models. This approach frames the problem as a knapsack-constrained subset…

  15. RESEARCH · CL_08628 ·

    New research proposes reasoning-aware training for better dialogue summarization

    Researchers have developed a new framework for multi-role dialogue summarization that moves beyond traditional overlap metrics like ROUGE. Their approach incorporates explicit cognitive-style reasoning and reward-based …

  16. RESEARCH · CL_04682 ·

    Eugene Yan explores challenges in evaluating abstractive summaries and detecting hallucinations

    Evaluating abstractive summarization, which involves rephrasing source material rather than copying sentences, presents challenges, particularly in assessing relevance and factual consistency. While fluency and coherenc…