PulseAugur
EN
LIVE 15:43:50
ENTITY LlamaGuard

LlamaGuard

PulseAugur coverage of LlamaGuard — every cluster mentioning LlamaGuard across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
2
6 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
2 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/1 · 6 TOTAL
  1. SIGNIFICANT · CL_186551 ·

    Mistral AI releases Shieldstral, a flexible 3B guard model

    Mistral AI has released Shieldstral, a 3 billion parameter guard model designed to classify content policy violations. Unlike previous models like LlamaGuard and ShieldGemma, Shieldstral's policy is embedded within the …

  2. TOOL · CL_160823 ·

    New TopoGuard defense tackles split-knowledge attacks in RAG systems

    A new defense mechanism called TopoGuard has been developed to combat split-knowledge attacks targeting retrieval-augmented generation (RAG) systems. These attacks involve injecting seemingly benign documents that, when…

  3. RESEARCH · CL_88573 ·

    Google's AMS tool finds critical safety flaws in three tested LLMs

    Google Cloud has open-sourced AMS (Activation Model Scanner), a tool that analyzes the geometric structure of a model's activation space to verify safety training. Unlike traditional behavioral tests, AMS directly inspe…

  4. TOOL · CL_57758 ·

    LLM Agents Vulnerable to Tool-Output Injection Attacks

    LLM agents possess a significant security vulnerability where malicious code can be injected through the outputs of tools they utilize. This 'tool-output injection' bypasses standard input and output guardrails because …

  5. RESEARCH · CL_16158 ·

    AI safety models vulnerable to fine-tuning and embedding bypass attacks

    Two new research papers explore vulnerabilities in AI safety mechanisms. The first paper, "When Safety Geometry Collapses," demonstrates how fine-tuning even benign guard models can inadvertently destroy their safety al…

  6. TOOL · CL_09472 ·

    New proxy tool blocks prompt injection attacks on AI models

    A new tool called Arc Gate has been developed to act as a proxy, sitting in front of any OpenAI-compatible endpoint. This proxy is designed to effectively block prompt injection attacks before they can reach the underly…