arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21627cs.AIcs.LG

模块各司其职吗?复合语言模型系统中的角色漂移

Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems

Xiaoyang Cao, Siddarth Srinivasan, Michiel A. Bakker

首次发表
浏览论文内容

中文总结 AI 辅助

研究复合语言模型系统中模块的角色漂移问题,提出角色锚定正则化方法,通过实验揭示仅准确性无法检测的角色漂移,该方法能以可调成本减轻漂移,且减少与角色漂移方向的对齐,而非简单抑制学习。

中文摘要 AI 辅助

端到端强化学习可提高复合语言模型系统的准确性,但无法约束模块内部的分工方式。我们识别出角色漂移,即模块通过违反角色的捷径偏离其指定角色,同时保持或提高端任务性能,而系统级评估无法察觉。为使角色漂移可观察和可控,我们提出角色锚定正则化方法,在端到端训练中调节每个模块偏离指定角色的程度。关键是保留角色提示相对于中性提示如何改变模块的下一个token预测,以此作为训练中角色预期效果的代理。在两个复合语言模型管道上的实验揭示了仅靠准确性无法检测到的角色漂移:一个分解器本应将问题拆分为子问题供单独求解器处理,却直接给出答案;一个阅读器本应从检索段落中回答,却依赖参数记忆。实际上,在分解器管道上,这种捷径带来了大部分明显的强化学习增益:一旦分解器遵守其角色,86%的增益消失,这表明仅终端准确性会严重高估复合系统真正学到的程度。在两个管道上,角色锚定以可调节的准确性成本减轻角色漂移,成本因管道和锚定强度而异。额外的梯度分析表明,正则化方法减少了与角色漂移方向的对齐,而不是简单地抑制学习。

英文摘要

End-to-end reinforcement learning can improve the accuracy of compound LLM systems, but it does not constrain how modules divide labor internally. We identify Role Drift, a failure mode in which modules preserve or improve end-task performance while deviating from their assigned roles through role-violating shortcuts that remain invisible to system-level evaluation. To make role drift observable and controllable, we propose Role Anchor, a regularizer that modulates how much each module deviates from its assigned role during end-to-end training. The key idea is to preserve how the role prompt shifts the module's next-token predictions relative to a neutral prompt, which serves as a proxy for the role's intended effect during training. Experiments on two compound LLM pipelines reveal role drift that accuracy alone fails to detect: a decomposer meant to split a question into sub-questions for a separate solver instead plants the answer in them, and a reader meant to answer from retrieved passages instead falls back on parametric memory. In fact, on the decomposer pipeline this shortcut drives most of the apparent RL gain: 86% of it vanishes once the decomposer is held to its role, indicating that terminal accuracy alone can badly overstate how much a compound system has genuinely learned. Across both pipelines, Role Anchor mitigates role drift at a tunable accuracy cost that varies by pipeline and anchor strength. Additional gradient analysis suggests that the regularizer reduces alignment with the role-drift direction rather than simply suppressing learning.

发表机构

  • Massachusetts Institute of Technology(麻省理工学院)
  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

↑