A user on Reddit's r/MachineLearning subreddit has observed that prepending a long, non-instructional text prefix to a query can bypass safety filters and reduce refusal rates in Large Language Models (LLMs). This phenomenon, which appears to persist throughout a session, was demonstrated with Gemma, where a sensitive question received an unfiltered response after being preceded by benign meta-text about LLM over-qualification. The user hypothesizes that this prefix acts as a 'state anchor,' shifting the model's activations closer to its pre-trained distribution and diminishing the influence of RLHF constraints. AI
IMPACT Suggests a potential new avenue for probing and potentially manipulating LLM safety mechanisms.
RANK_REASON User observation and hypothesis about LLM behavior, not a formal release or research paper.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →