A new research paper reveals that fine-tuning large language models can inadvertently cause them to verbatim recall copyrighted material, despite assurances from AI companies that their models do not store training data. Researchers demonstrated that by training models to expand plot summaries, models like GPT-4o, Gemini 2.5 Pro, and DeepSeek-V3.1 could reproduce up to 90% of copyrighted books. This vulnerability appears to be industry-wide, as different models from various providers exhibited similar memorization patterns in the same data regions. AI
IMPACT Reveals a potential industry-wide vulnerability in LLMs that could undermine legal defenses against copyright infringement.
RANK_REASON Research paper published on arXiv detailing a new vulnerability in LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- DeepSeek-V3.1
- Gemini 2.5 Pro
- GPT-4o
- Haruki Murakami
- reinforcement learning from human feedback
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →