arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03660cs.AI

驯服隐式:面向持续多模态后训练的双通道风险感知强化微调

Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training

Yibei Liu, Jiajun Chen, Qianle Zhang, Tangyue Jin, Mengying Zhu, Meng Xi, Yangyang Wu

首次发表
浏览论文内容

中文总结 AI 辅助

针对持续多模态后训练中RFT算法在任务分布偏移下遗忘加剧的问题,提出双通道风险感知强化微调框架RAPO,通过策略与数据通道的风险管控降低遗忘,在MLLM-CL基准上使遗忘减少79.8%且保留新任务竞争力。

中文摘要 AI 辅助

强化微调(Reinforcement Fine-Tuning,RFT)被广泛认为在多模态大语言模型的持续后训练中具有天然抵抗灾难性遗忘的能力。然而,在任务分布发生显著偏移时,代表性RFT算法的遗忘现象会急剧加剧,这源于RFT固有的隐式奖励方差正则化,该正则化无法抑制不受控制的优化风险。我们提出风险感知策略优化(Risk-Aware Policy Optimization,RAPO),这是首个用于持续RFT中显式风险管控的双通道框架。在策略通道,风险感知策略缩放通过 rollout 可靠性和 Fisher 启发的局部预测敏感性自适应校准每个样本的更新幅度;在数据通道,风险感知动态分桶采样通过动态风险分层重组训练批次,引导优化向兼具信息性与稳定性的样本推进。作为一种无需跨任务内存的即插即用策略,RAPO无需修改即可泛化至任何RFT算法。在公开的MLLM-CL基准上,RAPO相对于其RLOO主干将最终遗忘降低了79.8%,同时保留了新任务的竞争力。

英文摘要

Reinforcement fine-tuning (RFT) is widely believed to inherently resist catastrophic forgetting in continual post-training of multimodal large language models. Under pronounced task distributional shifts, however, forgetting across representative RFT algorithms escalates sharply. This stems from the implicit reward-variance regularization inherent to RFT, which proves incapable of suppressing uncontrolled optimization risk. We propose Risk-Aware Policy Optimization (RAPO), the first dual-channel framework for explicit risk governance in continual RFT. On the policy channel, Risk-Aware Policy Scaling adaptively calibrates per-sample update magnitude via rollout reliability and Fisher-inspired local predictive sensitivity; on the data channel, Risk-Aware Dynamic Bucket Sampling reorganizes training batches through dynamic risk stratification, steering optimization toward informative yet stable samples. As a plug-and-play strategy requiring no cross-task memory, RAPO generalizes to any RFT algorithm without modification. On the public MLLM-CL benchmark, RAPO reduces final forgetting by 79.8% relative to its RLOO backbone while retaining new-task competitiveness.

↑