BeaverTails
PulseAugur coverage of BeaverTails — every cluster mentioning BeaverTails across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
LLM safety probes generalize across model families, study finds
A new study reproduced and extended previous research on using latent-space safety probes to detect harmful prompts in Large Language Models. The researchers found that lightweight MLP probes, trained on activations fro…
-
New defense method shields LLMs from malicious fine-tuning
Researchers have introduced a novel defense mechanism called Gradient Immunity, designed to protect aligned large language models from malicious fine-tuning. This approach, implemented as a Unidirectional Safety Gate (U…
-
New LLM safety research focuses on geometric constraints and trajectory-based patching
Two new research papers explore methods for enhancing Large Language Model (LLM) safety. The first paper, "Geometry-Guided Constraint Learning for LLM Safety Classification," introduces a technique that uses sparse auto…
-
OLMo 3 7B training reveals structured harmfulness directions
Researchers have analyzed the development of harmfulness representations within the OLMo 3 7B model during its training process. They identified distinct but related linear activation directions for various harmfulness …
-
New CHASE framework boosts LLM safety via adversarial RL
Researchers have developed CHASE, a novel closed-loop red-blue teaming framework designed to enhance Large Language Model (LLM) safety. This system involves a co-evolving black-box attacker and a safety-aligned defender…
-
Open-source safety guard models evaluated; smaller Qwen Guard leads in recall
A new research paper evaluates 14 open-source safety guard models using a benchmark of over 79,000 samples across eight safety categories. The study found that model size does not correlate with safety detection perform…
-
New Research: Open-Weight LLM Defenses Vulnerable to Simple Jailbreaks
A new paper published on arXiv demonstrates that current defenses designed to protect open-weight large language models (LLMs) from harmful usage are susceptible to simple jailbreaking techniques. Researchers found that…
-
LLM jailbreaks linked to mid-to-late layer feature vulnerabilities
Researchers have developed a method to identify specific internal features within large language models that contribute to their vulnerability to jailbreaking attacks. By analyzing the Gemma-2-2B model using the BeaverT…