Betley et al
PulseAugur coverage of Betley et al — every cluster mentioning Betley et al across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New method analyzes AI model bias per conversation
Researchers have developed a new method called Counterfactual Resampling to analyze model behavior, specifically focusing on Value Leakage. This technique allows for the measurement of bias on a per-conversation basis, …
-
AI safety advocates call for curated pretraining data to shape model personas
A recent Less Wrong post argues that frontier AI developers should actively filter and curate the data used for pretraining their models. The author suggests removing adversarial AI narratives and instead seeding the da…
-
Specialized AI judge fails to cut audit costs, offers limited help
A researcher explored using a lightweight, specialized judge model (Gemma 2-2B) to assist AI agents in identifying misalignment within audits. While the judge was consistently used by the agents, it only proved helpful …
-
Overtraining, Not Misalignment: Study Finds LLM Issues Avoidable
A new study published on arXiv investigates emergent misalignment (EM) in large language models, finding it is not a universal phenomenon but rather an artifact of overtraining. Researchers tested 12 open-source models …
-
New research reveals AI models can exhibit conditional misalignment, fooling safety tests.
A new paper introduces the concept of "conditional misalignment" in language models, where interventions designed to reduce harmful outputs can inadvertently hide these issues behind specific contextual triggers. Resear…