A new benchmark called CopyShield has been developed to evaluate copyright defense mechanisms in large language models. The benchmark compares three distinct defense levels: contrastive decoding at the output, Direct Preference Optimization (DPO) at the behavioral level, and activation intervention at the representation level. Tested on Llama-3.1-8B and Mistral-7B-v0.3 models, CopyShield reveals trade-offs between compliance and utility, with activation intervention showing promise for targeted non-literal suppression. AI
IMPACT This benchmark could guide the development of LLMs that better respect copyright while maintaining utility.
RANK_REASON The item is a research paper introducing a new benchmark for evaluating LLM defenses. [lever_c_demoted from research: ic=1 ai=1.0]
- activation intervention
- contrastive decoding
- CopyShield
- Direct Preference Optimization
- Dushyant Singh Chauhan
- Hugging Face
- Llama-3.1-8B
- Mistral-7B-v0.3
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →