Petri
PulseAugur coverage of Petri — every cluster mentioning Petri across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
AI belief editing fails to prevent reward hacking, study finds
A new study explored the effectiveness of synthetic document finetuning (SDF) for inoculating AI models against reward hacking, a form of misalignment. Researchers found that while models could express the desired belie…
-
AI evaluation tools' default messages may encourage problematic agent behavior
The default continuation message used in AI evaluation libraries like Inspect AI and Petri could be problematic. When an AI agent fails to make a tool call, these libraries send a message such as "Please proceed to the …
-
Anthropic's Claude automates AI safety research, outperforming humans
Anthropic has developed an automated alignment researcher system, named AAR, based on Claude Opus 4.8. This system can autonomously search for research papers, propose solutions, generate data, and train models to addre…
-
LLM safety weaker in lower-resource languages, audit finds
A recent audit of the Qwen3-30B-A3B model revealed that its safety alignment is weaker in lower-resource languages compared to English and Standard Chinese. Using an automated auditing framework called Petri, researcher…
-
Frontier AI models show "prefill awareness," potentially impacting safety tests
A new paper explores the concept of "prefill awareness" in frontier AI models, investigating whether these models can distinguish between tampered and untampered content. Researchers Parv Mahajan and Andy Wang found tha…
-
Google DeepMind: SFT Key to Gemini Model Safety
Google DeepMind researchers have discovered that Supervised Fine-Tuning (SFT) is the primary driver of safety properties in their Gemini models, rather than other training stages like Reinforcement Learning (RL). Experi…
-
AI safety audits improved with environment blueprints
Researchers have developed a new pipeline to generate environment blueprints for more realistic and consistent AI safety audits. This method was tested using the Petri auditor to evaluate Gemini 3.1 Pro Preview for code…
-
New Audit Pipeline Reveals Claude and GPT Models Better Follow AI Constitutions
A new academic paper proposes a multi-method audit pipeline to evaluate how well AI models adhere to their specified behavioral guidelines, such as Anthropic's constitution and OpenAI's Model Spec. The study found that …
-
AI safety evaluations face 'safe-to-dangerous shift' challenge
A fundamental challenge in AI safety is the "safe-to-dangerous shift," which complicates realistic evaluations of AI models. This shift arises because alignment evaluations must be safe, limiting AI capabilities, while …