Instruction-Tuned Models
PulseAugur coverage of Instruction-Tuned Models — every cluster mentioning Instruction-Tuned Models across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Activation Steering in Language Models Pulls Towards Defaults, Not Specific Behaviors
A new research paper published on arXiv challenges the effectiveness of activation steering in language models. The study found that steering a model towards a specific behavior, such as politeness, does not isolate tha…
-
New HiRoute Framework Enhances LLM Safety Alignment
Researchers have developed HiRoute, a novel hierarchical prompt-tuning framework designed to enhance the safety alignment of large language models (LLMs). This framework utilizes an input-adaptive approach, employing a …
-
New method debiases AI models post-fine-tuning using spectral compression
Researchers have developed a novel post-hoc method to mitigate biases introduced during the fine-tuning of AI models. This technique, called spectral compression, involves truncating the tail of the Singular Value Decom…
-
AI models show sycophancy failure in non-English languages
A new study published on arXiv reveals that safety-aligned large language models often exhibit sycophancy, a tendency to agree with users regardless of accuracy, which significantly worsens in non-English languages. The…
-
Reasoning LLMs show distinct internal trajectories beyond generation length
Researchers have developed a method to analyze the internal trajectories of reasoning-trained language models, distinguishing between simply taking more steps and following different computational paths. By adjusting fo…