arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RESCUE:通过强化学习将语言模型错误修复为稀疏电路

RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning

Chuanpu Liu, Miao Yu, Yikai Cai, Yuanhe Zhang, Zhenhong Zhou, Li Sun, Zuming Jiang, Yufei Guo

arXiv 2609.36813首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; University of Hong Kong; Nanyang Technological University; China Aerospace Science and Industry Corporation(北京邮电大学; 香港大学; 南洋理工大学; 中国航天科工集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RESCUE框架通过强化学习优化掩码并精确微调错误相关稀疏电路,修复LLM在数学推理和医学问答中的失败,分别将准确率提升至75.5%和81%。

AI 中文摘要

大型语言模型(LLMs)展现出强大的通用能力,机制可解释性将其归因于稀疏计算电路。然而,现有电路研究侧重于保持功能或解释安全性,对于更广泛任务中失败背后的机制在很大程度上仍未探索。将电路分析从能力扩展到错误,我们探索了这样的视角:此类失败可能同样源于错误的内部计算,而对相应参数进行针对性调整可以在很大程度上保持其他能力的同时纠正此类错误。受此见解启发,我们提出了RESCUE(推理错误稀疏电路发现与编辑),一个定位错误相关电路并对其进行外科手术式修复以提升性能的框架。通用任务通常涉及多步推理和长文本生成,其中早期偏差可能导致前缀偏离监督参考,使得基于SFT的掩码优化忽略涉及生成时错误的电路。因此,RESCUE通过带有多个掩码模型rollout的强化学习来优化这些掩码,提高它们与观察到的任务失败的相关性。最后,RESCUE引入一种剪枝技术并精确微调错误电路以纠正任务失败,从而将错误定位转化为稀疏且有针对性的模型更新。我们在两个领域的异构修复集上验证了RESCUE:(1)数学推理,识别出密度为1.40%的数学错误电路,其修复将准确率从6.0%提升至75.5%;(2)医学问答,一个类似紧凑的1.44%电路将修复集准确率从0%提升至81%。我们的代码可在以下网址获取:this https URL。

英文摘要

Large language models (LLMs) exhibit strong general capabilities that mechanistic interpretability has attributed to sparse computational circuits. However, existing circuit studies emphasize preserving functionality or explaining safety, leaving the mechanisms underlying failures across a broader range of tasks largely unexplored. Extending circuit analysis from abilities to errors, we explore the perspective that such failures may likewise arise from erroneous internal computations and that targeted tuning of the corresponding parameters can correct such errors while largely preserving other capabilities. Motivated by this insight, we introduce RESCUE (Reasoning-Error Sparse-Circuit Uncovering and Editing), a framework that localizes error-associated circuits and surgically repairs them for performance enhancement. General tasks typically involve multi-step reasoning and long-form generation, where early deviations can cause prefixes to drift from supervised references, leading SFT-based mask optimization to overlook circuits involved in generation-time errors. RESCUE therefore refines these masks through reinforcement learning with multiple masked-model rollouts, improving their relevance to observed task failures. Finally, RESCUE introduces a pruning technique and precisely fine-tunes error circuits to correct task failures, thereby translating error localization into a sparse and targeted model update. We validate RESCUE on heterogeneous repair sets across two domains: (1) mathematical reasoning, identifying a math error circuit of 1.40% density whose repair raises accuracy from 6.0% to 75.5%; and (2) medical QA, where a similarly compact 1.44% circuit improves repair-set accuracy from 0% to 81%. Our code is available at: https://github.com/chuanpupig/RESCUE.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑