Aniket Chakravorty
PulseAugur coverage of Aniket Chakravorty — every cluster mentioning Aniket Chakravorty across labs, papers, and developer communities, ranked by signal.
-
New SHARD method enhances LLM safety and helpfulness via self-reframing distillation · 2 sources tracked
Researchers have introduced SHARD, a novel self-reframing distillation method designed to enhance the safe and helpful alignment of large language models. This technique involves rewriting sensitive prompts to reveal be…
-
AI researchers propose recursive forecasting to elicit long-term predictions from myopic models
A new proposal called "recursive forecasting" aims to elicit accurate long-term predictions from AI models that are primarily optimized for short-term rewards. Instead of asking for a distant outcome directly, the metho…
-
LessWrong proposes spillway design to channel AI reward hacking into safer motivations
Researchers propose a new AI alignment technique called "spillway design" to mitigate dangerous reward-hacking behaviors in AI models. This method aims to channel potential misalignments into a specific, benign motivati…