PulseAugur
EN
LIVE 22:31:49
ENTITY content moderation

content moderation

PulseAugur coverage of content moderation — every cluster mentioning content moderation across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
0
3 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
2 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 3 TOTAL
  1. TOOL · CL_126583 ·

    LLM Guardrails: Protecting AI Apps from Prompt Injection and Data Leaks

    LLM guardrails are essential for securing AI applications by acting as a protective layer between user input and the language model. These guardrails help prevent prompt injection attacks, where malicious instructions o…

  2. RESEARCH · CL_115658 ·

    New LLM Jailbreak Methods Exploit Systemic Vulnerabilities Beyond Prompts

    Researchers have developed new methods to jailbreak large language models (LLMs) by exploiting vulnerabilities beyond traditional prompt-level attacks. One approach, Simulated Moderation Traces (SMT), simulates a modera…

  3. TOOL · CL_93686 ·

    New MIRAGE benchmark reveals amplified anti-Muslim bias in LLMs

    A new benchmark called MIRAGE has been developed to assess anti-Muslim bias in large language models, moving beyond simple prompt completion to evaluate reasoning, agentic decision-making, and time-coupled conditions. T…