发表机构
The University of Queensland; Griffith University(昆士兰大学; 格里菲斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM推荐系统推理冗长导致成本高的问题,提出首个用于推理压缩的注意力头级细粒度模型合并框架,在不降低推荐准确率的前提下最多缩短24.3%的推理长度。
AI 中文摘要
基于大语言模型的推荐系统越来越多地采用“慢思考”模型,该模型在做出预测前会生成分步推理过程,通常比直接预测的“快思考”模型准确率更高。然而,这类模型的推理轨迹往往过于冗长,增加了推理成本却未带来相应的准确率提升。现有的基于训练的推理压缩方法通常会产生大量适配成本,而推理时方法则较为脆弱且难以扩展。这些局限性促使模型合并成为在共享参数空间中传递模型间专门行为的有前景的无训练方向。特别是,将慢思考模型与快思考模型合并,提供了一种平衡推荐准确率与推理简洁性的自然机制。为此,我们提出了据我们所知首个用于推荐系统推理压缩的模型合并框架。与在模型组件上应用统一合并系数的传统合并方法不同,我们的方法在单个注意力头层面执行细粒度合并,捕捉推荐推理中的异质模式。每个注意力头根据其对关键推理证据的贡献以及对参数变化的敏感性被分配不同的合并系数,从而实现将快思考模型的简洁行为选择性注入慢思考模型,在不损害推荐质量的前提下减少推理冗长性。在三个基准数据集上的实验表明,我们的方法可将推理长度最多降低24.3%,同时在维持推荐准确率方面优于有竞争力的模型合并基线。代码可在该https URL获取。
英文摘要
Large language model-based recommender systems are increasingly adopting slow-thinking models that generate step-by-step reasoning before making predictions, often achieving higher accuracy than fast-thinking models that predict directly. However, their reasoning traces are often unnecessarily verbose, increasing inference costs without commensurate accuracy gains. Existing training-based approaches to reasoning compression often incur substantial adaptation costs, while inference-time methods are brittle and difficult to scale. These limitations motivate model merging as a promising training-free direction for transferring specialised behaviours between models in a shared parameter space. In particular, merging a slow-thinking model with a fast-thinking counterpart provides a natural mechanism for balancing recommendation accuracy and reasoning conciseness. To this end, we propose, to our knowledge, the first model merging framework for reasoning compression in recommender systems. Unlike conventional merging methods that apply uniform merge coefficients across model components, our method performs fine-grained merging at the level of individual attention heads, capturing heterogeneous patterns in recommendation reasoning. Each attention head is assigned a distinct merge coefficient according to its contribution to critical reasoning evidence and its sensitivity to parameter change, enabling selective injection of the concise behaviour of the fast-thinking model into the slow-thinking model and reducing reasoning verbosity without compromising recommendation quality. Experiments on three benchmark datasets show that our method reduces reasoning length by up to 24.3% while outperforming competitive model merging baselines in maintaining recommendation accuracy. The code is available at https://github.com/linhledieu/REAM.