PulseAugur
EN
LIVE 08:53:10

Qwen3 LLM preferences for time-based decisions are steerable, study finds

Researchers have identified and manipulated temporal preferences within the Qwen3-32B large language model. By training contrastive linear probes, they discovered directions in the model's residual stream that represent short-term versus long-term decision-making. Applying contrastive activation-addition steering using these directions significantly altered the model's choices in temporal-choice tasks, including monetary intertemporal choices and a planning capability benchmark. This work demonstrates that LLM intertemporal preferences are measurable and steerable, with implications for AI systems providing advice on delayed costs and benefits, and for AI safety concerning long-horizon planning. AI

IMPACT Demonstrates steerability of LLM intertemporal preferences, impacting AI systems that advise on delayed costs/benefits and AI safety regarding long-horizon planning.

RANK_REASON Academic paper detailing a new method for steering LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3 LLM preferences for time-based decisions are steerable, study finds

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Michal Mr\'az, Justin Shenk ·

    Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition

    arXiv:2608.03892v1 Announce Type: new Abstract: We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-…