PulseAugur
EN
LIVE 15:25:14
ENTITY BeaverTails

BeaverTails

PulseAugur coverage of BeaverTails — every cluster mentioning BeaverTails across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
2
8 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
8 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/1 · 8 TOTAL
  1. TOOL · CL_193471 ·

    LLM safety probes generalize across model families, study finds

    A new study reproduced and extended previous research on using latent-space safety probes to detect harmful prompts in Large Language Models. The researchers found that lightweight MLP probes, trained on activations fro…

  2. TOOL · CL_185304 ·

    New defense method shields LLMs from malicious fine-tuning

    Researchers have introduced a novel defense mechanism called Gradient Immunity, designed to protect aligned large language models from malicious fine-tuning. This approach, implemented as a Unidirectional Safety Gate (U…

  3. RESEARCH · CL_154126 ·

    New LLM safety research focuses on geometric constraints and trajectory-based patching

    Two new research papers explore methods for enhancing Large Language Model (LLM) safety. The first paper, "Geometry-Guided Constraint Learning for LLM Safety Classification," introduces a technique that uses sparse auto…

  4. TOOL · CL_82169 ·

    OLMo 3 7B training reveals structured harmfulness directions

    Researchers have analyzed the development of harmfulness representations within the OLMo 3 7B model during its training process. They identified distinct but related linear activation directions for various harmfulness …

  5. TOOL · CL_72641 ·

    New CHASE framework boosts LLM safety via adversarial RL

    Researchers have developed CHASE, a novel closed-loop red-blue teaming framework designed to enhance Large Language Model (LLM) safety. This system involves a co-evolving black-box attacker and a safety-aligned defender…

  6. TOOL · CL_58669 ·

    Open-source safety guard models evaluated; smaller Qwen Guard leads in recall

    A new research paper evaluates 14 open-source safety guard models using a benchmark of over 79,000 samples across eight safety categories. The study found that model size does not correlate with safety detection perform…

  7. TOOL · CL_53861 ·

    New Research: Open-Weight LLM Defenses Vulnerable to Simple Jailbreaks

    A new paper published on arXiv demonstrates that current defenses designed to protect open-weight large language models (LLMs) from harmful usage are susceptible to simple jailbreaking techniques. Researchers found that…

  8. RESEARCH · CL_06616 ·

    LLM jailbreaks linked to mid-to-late layer feature vulnerabilities

    Researchers have developed a method to identify specific internal features within large language models that contribute to their vulnerability to jailbreaking attacks. By analyzing the Gemma-2-2B model using the BeaverT…