PulseAugur
EN
LIVE 11:42:28

New defense framework targets data poisoning in text summarization models

Researchers have developed a new framework to defend text summarization models against data poisoning attacks that occur during the fine-tuning stage. This method, called Detect, Unlearn, Restore, can identify poisoned data by analyzing training influence and behavioral sensitivity. The framework also includes a gradient-ascent unlearning technique to restore the model's original behavior with minimal loss in utility. AI

IMPACT This research offers a practical method to secure summarization models against malicious data manipulation, enhancing the reliability of AI-generated text.

RANK_REASON The cluster contains a research paper detailing a new defense framework for LLMs.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New defense framework targets data poisoning in text summarization models

COVERAGE [2]

  1. arXiv cs.CL TIER_1 English(EN) · Poojitha Thota, Shirin Nilizadeh ·

    Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning

    arXiv:2606.26036v1 Announce Type: new Abstract: Training-time data poisoning during fine-tuning poses a significant threat to large language models (LLMs) deployed for abstractive text summarization, where small task-specific datasets exert disproportionate influence on model beh…

  2. arXiv cs.CL TIER_1 English(EN) · Shirin Nilizadeh ·

    Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning

    Training-time data poisoning during fine-tuning poses a significant threat to large language models (LLMs) deployed for abstractive text summarization, where small task-specific datasets exert disproportionate influence on model behavior. In this setting, adversaries manipulate f…