Researchers have identified and manipulated temporal preferences within the Qwen3-32B large language model. By training contrastive linear probes, they discovered directions in the model's residual stream that represent short-term versus long-term decision-making. Applying contrastive activation-addition steering using these directions significantly altered the model's choices in temporal-choice tasks, including monetary intertemporal choices and a planning capability benchmark. This work demonstrates that LLM intertemporal preferences are measurable and steerable, with implications for AI systems providing advice on delayed costs and benefits, and for AI safety concerning long-horizon planning. AI
IMPACT Demonstrates steerability of LLM intertemporal preferences, impacting AI systems that advise on delayed costs/benefits and AI safety regarding long-horizon planning.
RANK_REASON Academic paper detailing a new method for steering LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →