ENTITY
Adversarial prompts
Adversarial prompts
PulseAugur coverage of Adversarial prompts — every cluster mentioning Adversarial prompts across labs, papers, and developer communities, ranked by signal.
Total · 30d
0
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
2 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 2 TOTAL
-
New framework uses computation graphs to diagnose LLM jailbreak vulnerabilities
Researchers have developed a new framework to understand how large language models (LLMs) are vulnerable to adversarial prompts and jailbreak attacks. This method uses paired internal computation graphs to represent pro…
-
New research tackles multilingual LLM toxicity detection and mitigation
Two new research papers explore methods for detecting and mitigating toxicity in large language models (LLMs), particularly focusing on multilingual contexts. The first paper surveys existing strategies for identifying …