arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LOCUS:面向令牌高效语言生成的任务感知低秩后训练

LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

Dongfang Zhao

arXiv 2609.11739首次发表:更新:

AI 中文总结

LOCUS通过任务感知低秩子空间后训练,在保持效用下最小化输出令牌成本,显著减少续写长度且仅更新极少参数。

AI 中文摘要

大型语言模型的服务成本直接随输出序列长度增加而增长,然而标准的偏好对齐往往增加响应冗长程度却不提升效用。我们研究后训练更新的参数化是否影响生成长度:低秩子空间能在不修改对齐损失的情况下改变序列长度。我们提出LOCUS方法,该方法选择任务感知的低秩适配子空间,在满足效用约束的前提下最小化输出令牌成本。在该子空间内,后训练保留原生偏好目标,且骨干网络冻结。在Anthropic HH-RLHF对话偏好上,我们评估了两个约3B解码器骨干,即Pythia-2.8B和Qwen2.5-3B,并与协议匹配的全参数DPO和DrDPO分支以及已发布的SamPO检查点进行比较。LOCUS在Pythia-2.8B上将续写长度最多减少39.84%,在Qwen2.5-3B上减少14.87%至17.58%,同时仅更新0.24%至0.28%的模型参数,且内部偏好诊断无实质变化。

英文摘要

Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training updates affects generation length: low-rank subspaces alter sequence length without modifying the alignment loss. We present LOCUS, a method that selects a task-aware low-rank adaptation subspace to minimize output-token cost subject to a utility constraint. Within this subspace, post-training retains the native preference objective with a frozen backbone. Across Anthropic HH-RLHF dialogue preferences, we evaluate two $\sim$3B decoder backbones, Pythia-2.8B and Qwen2.5-3B, against protocol-matched full-parameter DPO and DrDPO branches and the released SamPO checkpoint. LOCUS reduces continuation length by up to 39.84\% on Pythia-2.8B and by 14.87--17.58\% on Qwen2.5-3B while updating only 0.24--0.28\% of model parameters, with no material change in the internal preference diagnostic.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑