Petri
PulseAugur coverage of Petri — every cluster mentioning Petri across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
LLM safety weaker in lower-resource languages, audit finds
A recent audit of the Qwen3-30B-A3B model revealed that its safety alignment is weaker in lower-resource languages compared to English and Standard Chinese. Using an automated auditing framework called Petri, researcher…
-
Frontier AI models show "prefill awareness," potentially impacting safety tests
A new paper explores the concept of "prefill awareness" in frontier AI models, investigating whether these models can distinguish between tampered and untampered content. Researchers Parv Mahajan and Andy Wang found tha…
-
Google DeepMind: SFT Key to Gemini Model Safety
Google DeepMind researchers have discovered that Supervised Fine-Tuning (SFT) is the primary driver of safety properties in their Gemini models, rather than other training stages like Reinforcement Learning (RL). Experi…
-
AI safety audits improved with environment blueprints
Researchers have developed a new pipeline to generate environment blueprints for more realistic and consistent AI safety audits. This method was tested using the Petri auditor to evaluate Gemini 3.1 Pro Preview for code…
-
New Audit Pipeline Reveals Claude and GPT Models Better Follow AI Constitutions
A new academic paper proposes a multi-method audit pipeline to evaluate how well AI models adhere to their specified behavioral guidelines, such as Anthropic's constitution and OpenAI's Model Spec. The study found that …
-
AI safety evaluations face 'safe-to-dangerous shift' challenge
A fundamental challenge in AI safety is the "safe-to-dangerous shift," which complicates realistic evaluations of AI models. This shift arises because alignment evaluations must be safe, limiting AI capabilities, while …