Bleu
PulseAugur coverage of Bleu — every cluster mentioning Bleu across labs, papers, and developer communities, ranked by signal.
8 day(s) with sentiment data
-
BLEU and ROUGE metrics explained for language model evaluation
BLEU and ROUGE are key metrics used to evaluate the performance of language models, particularly in tasks like machine translation and text summarization. BLEU focuses on precision of n-grams and includes a penalty for …
-
LLM-as-a-Judge: Using AI to Evaluate AI Output
The "LLM-as-a-Judge" technique utilizes a large language model to evaluate the output of other models, addressing the bottleneck of performance assessment in AI development. This method acts as a scalable and explainabl…
-
AI models win pun translation contest with novel multi-agent approach · 3 sources tracked
Researchers have developed a novel approach to translate puns from English to French, achieving first and second place in the CLEF JOKER 2025 Task 2 competition. The method combines large language models with specialize…
-
FAU researchers achieve top ranks in ImageCLEF 2026 multimodal reasoning task
Researchers from Florida Atlantic University (FAU) have developed a system for the ImageCLEF 2026 task on Multimodal Reasoning, focusing on Visual Multiple Choice Question Answering (Visual MCQ) and Visual Open Question…
-
Cross-lingual transfer in Turkic languages shows strong pair-specific performance
Researchers have investigated cross-lingual transfer techniques for machine translation within the Turkic language family, focusing on Turkish, Azerbaijani, Uzbek, Kazakh, and Kyrgyz. Their findings indicate that transf…
-
AI-generated counterspeech can be personalized for greater impact
Researchers have developed and evaluated strategies for generating contextualized AI-powered counterspeech to combat online toxicity. Unlike generic approaches, these new methods adapt to the conversational context and …
-
New research tackles personalized text generation and evaluation challenges
Two recent arXiv papers explore the challenges and methodologies of style-personalized text generation using large language models. The first paper, "PrefReward," introduces a new framework that explicitly models user p…
-
New benchmark PathReportEval standardizes pathology report generation evaluation
Researchers have introduced PathReportEval, a new benchmark and evaluation framework designed to standardize the assessment of pathology report generation from whole-slide images. This framework addresses the limitation…
-
Whisper model fine-tuned for robust Assamese speech recognition
Researchers have developed a fine-tuned version of the Whisper model to improve Automatic Speech Recognition (ASR) for the Assamese language. The fine-tuned model, trained on the Mozilla Common Voice 24.0-Assamese corpu…
-
New framework improves speech translation with targeted terminology adaptation
Researchers have developed a new framework called EGTA (Evidence-Grounded Terminology Adaptation) to improve simultaneous speech translation, particularly for technical content. EGTA focuses on adapting the translation …
-
New research paper calls for improved evaluation of personalized dialogue systems
A new research paper published on arXiv proposes a shift in how retrieval-augmented personalized dialogue systems are evaluated. The study highlights that current metrics like BLEU, ROUGE, and F1 fail to capture the dee…
-
LLMs fail to translate Korean Braille, study finds
A new research paper reveals significant accessibility failures in state-of-the-art Large Language Models (LLMs) when it comes to translating Korean Braille. Despite expectations that these models could handle Braille t…
-
CoPiT pipeline boosts low-resource Mongolian translation accuracy
Researchers have developed CoPiT, a novel translation pipeline designed to address the challenges of low-resource languages, specifically focusing on Mongolian. This system leverages the imbalance in data availability b…
-
Study: Prompt design boosts GPT-5.2 translation quality for journalists
A new study published on arXiv explores how prompt design affects the quality of Spanish-to-Chinese journalistic translations generated by GPT-5.2. Researchers tested 48 conditions, varying prompt types and languages, a…
-
New MBR decoding approach incorporates bidirectional effects for improved text generation
Researchers have introduced a novel noisy channel decomposition for Minimum Bayes Risk (MBR) decoding, aiming to improve text generation quality. This approach addresses the asymmetry in common evaluation metrics like B…
-
New GRAG framework enhances personalized conversational AI
Researchers have introduced GRAG, a new framework designed to improve personalized conversational systems, particularly in environments with limited resources or strict privacy requirements. GRAG decouples the complex t…
-
New AI models tackle low-resource Tangkhul-English translation
Researchers have developed two neural machine translation systems for the low-resource Tangkhul-English language pair. The primary system, utilizing ByT5-large fine-tuned on over 38,000 parallel sentences, achieved a BL…
-
LLMs struggle with Hausa and Fongbe translation, metrics unreliable
A new study evaluated the machine translation capabilities of four large language models (LLMs) for Hausa and Fongbe, two West African languages. The research found that while Hausa achieved acceptable translation quali…
-
New SPRI method enhances AI model upcycling under data constraints
Researchers have developed a new method called SVD-Partitioned Residual Initialization (SPRI) to improve the process of converting dense AI models into more efficient Mixture of Experts (MoE) models, a technique known a…
-
New methods advance simultaneous speech translation quality and evaluation
Researchers have developed new methods for evaluating and improving simultaneous speech translation systems, particularly for long-form content. One paper introduces a practical evaluation framework that measures senten…