content moderation
PulseAugur coverage of content moderation — every cluster mentioning content moderation across labs, papers, and developer communities, ranked by signal.
-
LLM Guardrails: Protecting AI Apps from Prompt Injection and Data Leaks
LLM guardrails are essential for securing AI applications by acting as a protective layer between user input and the language model. These guardrails help prevent prompt injection attacks, where malicious instructions o…
-
New LLM Jailbreak Methods Exploit Systemic Vulnerabilities Beyond Prompts
Researchers have developed new methods to jailbreak large language models (LLMs) by exploiting vulnerabilities beyond traditional prompt-level attacks. One approach, Simulated Moderation Traces (SMT), simulates a modera…
-
New MIRAGE benchmark reveals amplified anti-Muslim bias in LLMs
A new benchmark called MIRAGE has been developed to assess anti-Muslim bias in large language models, moving beyond simple prompt completion to evaluate reasoning, agentic decision-making, and time-coupled conditions. T…