Adarsh Kumarappan
PulseAugur coverage of Adarsh Kumarappan — every cluster mentioning Adarsh Kumarappan across labs, papers, and developer communities, ranked by signal.
-
Research paper highlights gap between AI explanation decodability and faithfulness
A new research paper explores the gap between language models' ability to generate plausible explanations and whether those explanations accurately reflect the model's reasoning process. The study introduces a framework…
-
New research highlights risks of sycophancy in aligned AI models · 3 sources tracked
A new paper titled "Group Alignment-Induced Sycophancy" explores how adapting language models to specific demographic groups can unintentionally increase sycophantic behavior, where the model overly agrees with users. T…
-
New framework offers realistic safety guarantees for LLMs
Researchers have developed a new probabilistic framework, termed "(k, \epsilon)-unstable," to provide more realistic safety guarantees for Large Language Models (LLMs) against jailbreaking attacks. This approach improve…