BERTScore: Evaluating text generation with BERT
PulseAugur coverage of BERTScore: Evaluating text generation with BERT — every cluster mentioning BERTScore: Evaluating text generation with BERT across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
-
AI legal assistants developed for Nepal to improve access to justice · 2 sources tracked
Two research papers introduce AI-powered legal assistants for Nepal, aiming to improve access to justice. The first, NepKANUN, utilizes a retrieval-augmented generation (RAG) framework with a fine-tuned large language m…
-
LoRA fine-tuning enhances Qwen2.5 models for control systems Q&A
Researchers have evaluated the effectiveness of LoRA fine-tuning on Qwen2.5 models for answering questions in a Linear Control Systems course. The study found that LoRA improved both textual similarity to reference answ…
-
LexFlip diagnostic tool evaluates legal meaning preservation in simplified text
Researchers have developed LexFlip, a new diagnostic tool designed to evaluate how well legal meaning is preserved when text is simplified. Current metrics often fail to distinguish between lexical overlap and actual le…
-
New PetQA benchmark evaluates AI veterinary knowledge
Researchers have developed PetQA, a new benchmark designed to evaluate the veterinary knowledge and clinical reasoning capabilities of large language models (LLMs) and large vision-language models (LVLMs). The benchmark…
-
LLMs show promise in extracting software design decisions from code commits
A preliminary study explored the ability of four Large Language Models (LLMs) to extract Architectural Design Decisions (ADDs) from source code commits. Models including Gemini 3-Pro, DeepSeek-R1, Kimi K2, and Qwen3 wer…
-
New corpus StageWell enhances AI support dialogues
Researchers have introduced StageWell, a new Chinese corpus designed for positive psychology support dialogues. This corpus, developed using the HQS protocol, organizes support into a six-stage process and includes 12,4…
-
PermitGPT uses generative AI for construction governance and safety
Researchers have developed PermitGPT, a generative AI framework designed to streamline urban construction governance. This system unifies scattered data from municipal and regulatory sources to identify safety hazards, …
-
SimpCue prompts boost multilingual text simplification modestly
Researchers have developed SimpCue, a novel prompting technique for multilingual text simplification using the Qwen3_8B language model. This method involves enriching prompts with explicit linguistic cues about sentence…
-
LLMs Compared for ASR Evaluation: Encoder vs. Generative Models
A new study published on arXiv explores the effectiveness of encoder and decoder-based Large Language Models (LLMs) for evaluating Automatic Speech Recognition (ASR) systems. The research compares metrics like BERTScore…
-
Gaze-supervised AI enhances chest X-ray diagnosis and report generation
Researchers have developed a novel two-stage multimodal framework for interpreting chest X-rays, integrating radiologist eye-tracking data to improve diagnostic accuracy and report generation. The first stage employs a …
-
New pipeline C-FEX improves factuality evaluation for domain-specific AI text generation
Researchers have developed a new evaluation pipeline called C-FEX to assess the factuality of generated text, particularly for domain-specific document generation tasks. This pipeline introduces Parametric Knowledge Pre…
-
LLMs show limited pedagogical competence in language teaching, study finds
A new research paper explores the pedagogical capabilities of large language models (LLMs) in language education. The study evaluated several state-of-the-art LLMs on their ability to identify, correct, and explain comm…
-
New metric 'Information Satisfaction' proposed for summarization evaluation
A new research paper proposes "Information Satisfaction" as a reader-centered metric for evaluating summarization systems. The authors argue that existing metrics like ROUGE and BERTScore, and even LLM-as-a-judge approa…
-
New 4B-parameter vision-language model targets 3D MRI analysis
Researchers have developed Mr3D-VL, a new 4-billion parameter vision-language foundation model specifically designed for multi-parametric 3D magnetic resonance imaging (mpMRI). This model addresses limitations in curren…
-
AI models win pun translation contest with novel multi-agent approach · 3 sources tracked
Researchers have developed a novel approach to translate puns from English to French, achieving first and second place in the CLEF JOKER 2025 Task 2 competition. The method combines large language models with specialize…
-
MIDAS framework enhances enterprise text summarization with multi-LLM adaptation
Researchers have introduced MIDAS (Multi-LLM Iterative Data-Adaptive Summarization), a novel framework designed to enhance text summarization for enterprise applications. MIDAS utilizes a multi-LLM approach that incorpo…
-
AI-generated counterspeech can be personalized for greater impact
Researchers have developed and evaluated strategies for generating contextualized AI-powered counterspeech to combat online toxicity. Unlike generic approaches, these new methods adapt to the conversational context and …
-
LLM framework verifies scientific paper claims against methods
Researchers have developed a new framework to enhance the peer review process for scientific papers by using large language models (LLMs) to verify claims within the papers themselves. This system, called intra-paper cl…
-
New LLM framework grounds ECG diagnosis in clinical knowledge
Researchers have developed a novel multimodal LLM framework designed to improve the explainability and trustworthiness of AI-driven cardiac diagnosis using electrocardiograms (ECGs). This new approach anchors report gen…
-
New RLHF framework improves Vietnamese translation of historical manuscripts
Researchers have developed a new multimodal Reinforcement Learning from Human Feedback (RLHF) framework to translate historical Han-Nom manuscripts into modern Vietnamese. This approach leverages both the visual informa…