Alex Mallen
PulseAugur coverage of Alex Mallen — every cluster mentioning Alex Mallen across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Analysis finds Anthropic's Mythos Preview alignment assessment has critical gaps
A recent analysis of Anthropic's Mythos Preview alignment risk assessment suggests that while the model may not exhibit unknown dangerous propensities, the assessment itself has significant limitations. The author argue…
-
New SHARD method enhances LLM safety and helpfulness via self-reframing distillation · 2 sources tracked
Researchers have introduced SHARD, a novel self-reframing distillation method designed to enhance the safe and helpful alignment of large language models. This technique involves rewriting sensitive prompts to reveal be…
-
AI insider info equals 2.5 months future knowledge
A researcher estimates that working inside a frontier AI company provides an informational advantage equivalent to having access to semi-public information about AI developments approximately 2.5 months into the future.…
-
AI model capabilities transfer less on difficult tasks
Researchers investigated how well AI model capabilities transfer across different behavioral tendencies, such as writing in bold versus plain text. They found that for simple tasks, capabilities transferred completely, …