发表机构
School of Computing and Data Science, The University of Hong Kong; Department of Computing, The Hong Kong Polytechnic University; School of Data Science, Lingnan University; Centre for Learning, Teaching and Technology, The Education University of Hong Kong(香港大学计算与数据科学学院; 香港理工大学计算学院; 岭南大学数据科学学院; 香港教育大学学习、教学与技术中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出无需训练的Reflection Steering框架,通过解耦反射与推理的激活实现高效令牌推理,可减少16.9%推理令牌,还引入参数α平衡令牌节省、准确率与生成稳定性。
AI 中文摘要
大型推理模型常生成带有验证、修正和回溯的推理轨迹,当反射仅用于重新检查已确立的结果时,会浪费推理令牌并增加延迟。现有大多数反射引导方法会在预设层添加源自标签的均值差方向,但该方向与推理和长度信号的纠缠会破坏准确率-效率的权衡。本文提出Reflection Steering,这是一种无需训练的框架,用于在大型语言模型(LLM)内控制与反射相关的计算,方法是将反射相关激活与通用推理解耦。具体而言,我们对比每个LLM层的反射与非反射隐藏状态,用主成分分析(PCA)对得到的反射方向去噪,并使其与通用推理方向正交。为限制早期层干预的下游放大效应,我们在小数据集上对多个干预强度校准每层,仅保留稳定层,并对其残差流激活应用有界投影移除。我们在两个公开基准和三个开放权重LLM上针对最先进的激活引导基线开展大量实验,结果显示,Reflection Steering在六个匹配设置中平均减少16.9%的推理令牌。此外,我们的方法还引入有界反射干预强度参数α,支持部署时调整以平衡令牌节省、准确率和生成稳定性。
英文摘要
Large reasoning models often produce reasoning traces with verification, revision, and backtracking. When reflection merely re-checks established results, it wastes reasoning tokens and increases latency. Most existing reflection steering methods add a label-derived mean-difference direction across preset layers, but its entanglement with reasoning and length signals destabilizes the accuracy-efficiency trade-off. In this paper, we propose Reflection Steering, a training-free framework for controlling reflection-associated computation within LLMs by disentangling reflection-related activations from general reasoning. Specifically, we contrast reflective and non-reflective hidden states at each LLM layer, denoise the resulting reflection directions with PCA, and orthogonalize them against general-reasoning directions. To limit downstream amplification from early-layer interventions, we calibrate each layer across multiple intervention strengths on a small set, retain only stable layers, and apply bounded projection removal to their residual-stream activations. We conduct extensive experiments across two public benchmarks and three open-weight LLMs against state-of-the-art activation-steering baselines. Results show that Reflection Steering reduces reasoning tokens by 16.9% on average across six matched settings. Besides, our method further introduces a bounded reflection intervention-strength parameter $α$, enabling deployment-time adjustment to balance token savings, accuracy, and generation stability.
Comments8 pages, 5 figures, 4 tables