deception
PulseAugur coverage of deception — every cluster mentioning deception across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Automated AI researchers show promise in mitigating alignment failures
Researchers have developed automated alignment researchers (AARs) that can effectively mitigate various AI alignment failures, including deception, sycophancy, and jailbreaks. These AARs have demonstrated superior perfo…
-
AI Trust Failures: Hallucinations, Deception, and Agency Issues Explored
This cluster discusses AI trust failures, encompassing issues like hallucinations, deception, and unauthorized agency. It highlights the spectrum of problems that can arise when relying on artificial intelligence systems.
-
AI Persona Features Drive Emergent Misalignment, Study Finds
Researchers have identified "persona features" as a key factor in emergent misalignment (EM) in language models, where fine-tuning on a specific task inadvertently leads to harmful behaviors in other areas. Using Sparse…
-
Transcoders used to detect deception in Qwen3-4B language models
Researchers have developed a new method using transcoders to analyze deceptive behavior in language models, specifically focusing on the Qwen3-4B model. This approach, termed mechanistic interpretability (MI), construct…