通过对比激活加法实现Qwen3的跨期偏好调控
Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition
浏览论文内容
中文总结 AI 辅助
本研究针对Qwen3-32B大模型,通过对比线性探针识别时间范围方向,利用对比激活加法调控实现跨期偏好的双向大幅改变,还提升了规划相关能力,为相关AI系统及安全问题提供支撑。
中文摘要 AI 辅助
我们研究了大语言模型Qwen3-32B中时间范围的线性表示,并利用这些表示改变模型与时间相关的偏好、推荐及能力。我们在教师强制的时间选择答案上训练对比线性探针,以在模型的残差流中找到短期与长期的方向,并在保留的二元时间选择任务、分布外的货币跨期选择任务以及TravelPlanner能力基准上评估对比激活加法调控。核心结果表明,可通过简单的对比线性探针识别时间范围方向,进而用于调控以诱导大幅双向偏好变化。在奖励规模和延迟均变化的分布外货币选择任务中,调控会强烈双向改变模型在较小-较近与较大-较晚奖励间的无差异阈值。我们还表明,在适度时间调控下,与规划相关的能力指标有所提升。这些结果表明,模型的跨期偏好是可测量且可调控的,这与涉及延迟成本和收益的AI系统建议,以及关于长时规划的安全问题相关。
英文摘要
We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-forced temporal-choice answers to find a short-term versus long-term direction in the model's residual stream, and evaluate contrastive activation-addition steering on a held-out binary temporal-choice task, an out-of-distribution monetary intertemporal-choice task, and a TravelPlanner capability benchmark. The central result is that temporal-horizon directions can be identified with simple contrastive linear probes and then used for steering to induce large, bidirectional preference changes. On an out-of-distribution monetary choice task that varies reward size and delay, steering strongly shifts the model's indifference threshold between smaller-sooner and larger-later rewards in both directions. We further show improvements on a planning-related capability metric under moderate temporal steering. These results suggest that model intertemporal preferences are measurable and steerable, which is relevant for AI systems that give advice involving delayed costs and benefits, and for safety questions about long-horizon planning.