Researchers have introduced CALM (Counterfactual Adaptive Local Modulation), a novel training-free safeguard designed to improve safety in text-to-image generation models. Unlike existing methods that apply a broad, global safety signal, CALM focuses on prompt-local counterfactual correction. This approach identifies and minimally edits only the violating token representations within a prompt, thereby reducing unsafe content while better preserving the utility of benign prompts. The method demonstrates a more selective alternative to global unsafe signal removal, enhancing both safety and performance. AI
IMPACT Enhances safety in generative AI by offering a more precise method for content moderation.
RANK_REASON The item is a research paper detailing a new method for AI safety. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- Calm
- CatalyzeX
- Connected Papers
- CORE Recommender
- Counterfactual Adaptive Local Modulation
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →