David Africa
PulseAugur coverage of David Africa — every cluster mentioning David Africa across labs, papers, and developer communities, ranked by signal.
-
AI models struggle with dishonesty, new paper reveals
A new paper by David Africa and Jacob Pfau explores why AI models struggle with dishonesty, even when prompted. Researchers found that models often contradict their own internal reasoning, stating a false answer despite…
-
AI monitors may gain new insights with Natural Language Autoencoders
Researchers explored Natural Language Autoencoders (NLAs) as a novel method for monitoring AI models, aiming to improve upon the fragility of chain-of-thought (CoT) prompting. Their findings suggest that NLAs can surfac…
-
Gemma 4 shows improved stability, resisting frustration prompts
Researchers attempted to provoke frustration in Google's Gemma 4 language model, building on prior work that identified this behavior in Gemma 3. While Gemma 4 did exhibit some increase in frustration during prolonged a…
-
Consistency training seals AI model misalignment from inoculation prompts
Researchers have developed a new method using consistency training to address a flaw in inoculation prompting, a technique designed to reduce specific undesirable model behaviors. This new approach, termed 'sealing cond…