Researchers have developed a new attack called Targeted Active Search (TAS) that can extract forgotten prompts from unlearned AI models. Unlike previous methods that assumed knowledge of the forgotten prompts, TAS uses retained data and black-box access to identify and reconstruct the prompts themselves. Experiments show TAS achieves 100% accuracy in identifying forgotten entities and reconstructs up to 95% of prompts with significantly fewer queries than naive probing methods. AI
IMPACT This research highlights a potential vulnerability in AI unlearning techniques, suggesting that even after data removal, sensitive information might still be recoverable.
RANK_REASON The cluster contains an academic paper detailing a new method for extracting information from AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- CORE Recommender
- DagsHub
- Direct Preference Optimization
- Hugging Face
- IArxiv Recommender
- LUNAR
- Targeted Active Search
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →