Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
PulseAugur coverage of Inference-Time Intervention: Eliciting Truthful Answers from a Language Model — every cluster mentioning Inference-Time Intervention: Eliciting Truthful Answers from a Language Model across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New Gated Activation Steering method combats LLM sycophancy and hallucination
Researchers have developed a new method called Gated Activation Steering to reduce sycophancy and hallucination in large language models, particularly for medical question answering. This technique uses Inference Time I…
-
New research questions effectiveness of activation steering in language models
A new research paper explores the phenomenon of activation steering in language models, questioning whether observed gains reflect intended control or compatibility with answer encodings. The study introduces Cross-Enco…
-
New method probes what activation steering truly controls in language models
Researchers have introduced a new evaluation method called Cross-Encoding Steering Evaluation to better understand what activation steering controls in language models. This method aims to distinguish between genuine co…
-
New research reveals "inverted" steering vectors in LLMs
Researchers have identified an "inverted detection-control phenomenon" in steering vectors (SVs), a technique used to influence the output of large language models. This phenomenon occurs when highly discriminative SVs,…
-
New methods tackle catastrophic forgetting in continual learning · 8 sources tracked
Researchers are developing new methods to address catastrophic forgetting in continual learning, a challenge where models lose previously acquired knowledge when learning new tasks. Several papers propose novel techniqu…