arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习决策,而非推理:通过低秩激活引导实现参数高效的决策算子

Learning to Decide, Not to Reason: Parameter-Efficient Decision Operators via Low-Rank Activation Steering

Ran Li, Lei Chen

arXiv 2610.06950首次发表:更新:

发表机构

Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种低秩激活引导的决策算子,通过行为克隆大幅降低技能注入成本,以少量参数匹配强化学习性能,并支持技能组合与迁移。

AI 中文摘要

向冻结的语言模型注入技能目前需要上百万参数和强化学习流程。我们引入了\method{},一种通过行为克隆训练的系统-1决策算子,将这一成本降低约两个数量级。默认算子使用330K参数即可匹配用强化学习训练的1.33M参数算子,达到或超过其性能,同时将3,685个token的深思熟虑压缩为6个token的决策,且精度无损。一个秩为4的变体仅用23K参数(最强已发布技能算子的1/58)即可胜任SearchQA,并接近胜任LiveMath(更高秩仍有帮助);同样的配方可迁移至五个任务和三个骨干模型,在训练后数月发布的LiveMath问题上仍保持分布外增益。与先前工作的差距在于可训练性,这由初始化和架构共同决定:先前算子的初始化在第一步优化时使两个大因子矩阵的梯度为零,而我们在共享低秩骨干内的零初始化输出投影能立即获得梯度,梯度流探针直接证实了这一点。增益并非思维链压缩:LiveMath的57个点中有23个超过了基础模型的最佳8次采样,且logit-lens探针显示算子沿模型现有的晚期层路径放大答案,而非更早写入。增益与基础模型在13个基础-任务对上的余量相关,技能可作为近似线性算子进行组合,可在推理时添加、插值并热切换。代码见此https URL。

英文摘要

Injecting skills into a frozen language model currently costs a million parameters and a reinforcement-learning pipeline. We introduce DecSteer, a System-1 decision operator trained by behavior cloning that lowers this cost by roughly two orders of magnitude. The default operator uses 330K parameters to match a 1.33M-parameter operator trained with reinforcement learning, exceeds or achieve comparable performance, while collapsing 3,685-token deliberation into a 6-token decision with no loss in accuracy. A rank-4 variant with 23K parameters, 1/58 of the strongest published skill operator, suffices for SearchQA and near-suffices for LiveMath, where higher rank still helps; the same recipe transfers across five tasks and three backbones, with out-of-distribution gains persisting on LiveMath problems released months after training. The gap to prior work is trainability, and it is set jointly by initialization and architecture. The initialization of prior operators zeroes the gradient of both large factor matrices at the first optimization step, whereas our zero-initialized output projection inside a shared low-rank backbone receives a gradient immediately, which a gradient-flow probe confirms directly. The gain isn't chain-of-thought compression. 23 of 57 LiveMath points beat the base model's best-of-8 sampling, and a logit-lens probe shows the operator amplifies the answer along the model's existing late-layer pathway, not writing it earlier. Gains track the base model's headroom across 13 base-task pairs, and skills compose as approximately linear operators that can be added, interpolated, and hot-swapped at inference time.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑