发表机构
Hong Kong Baptist University; City University of Hong Kong; University of Toronto; The University of Sydney; HKUST(香港浸会大学; 香港城市大学; 多伦多大学; 悉尼大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对强化学习引发的思维链混淆,提出TAME方法,利用稀疏自编码器定向抑制模板激活,在多个数据集上将可监控性提升至多30.9个百分点,兼顾任务准确率。
AI 中文摘要
强化学习(RL)提升了视觉语言模型(VLM)的推理能力,但可能引发思维链(CoT)混淆:这是一种操作性的、非故意的结果,即任务奖励或准确率上升,而推理痕迹变得不那么可接地和可监控。先前的工作主要从行为层面记录这种退化,其表征层面的相关性和可操作的控制手段尚不明确。我们发现,在RL过程中,模板相关和地面真值相关的激活变得不那么可分离;匹配的干预支持了所选特征对可监控性退化的贡献。在此证据指导下,我们提出了带有机制强制执行的目标反混淆方法(TAME),该方法利用稀疏自编码器(SAE)将行为反馈与RL期间对模板相关激活的定向抑制相结合。其非对称约束仅对高于其RL前基线的模板激活进行惩罚,锚定局部特征,同时行为反馈促进接地改进。在VIRL-39k、SPA-VL以及两个模型家族上,TAME相对于组相对策略优化(GRPO)将CoT可监控性分别提高了高达30.9和16.7个百分点。盲法人类评估发现两个数据集上的人类可监控性更高,两个保留的监控家族也复现了可监控性提升。任务准确率变化较小且混合,通用能力基准显示任务特定的权衡。这些结果为从行为监控到表征层面监督提供了路径,以实现更可审计的RL训练多模态系统。
英文摘要
Reinforcement learning (RL) improves reasoning in vision-language models (VLMs) but can induce chain-of-thought (CoT) obfuscation: an operational, non-intentional outcome where task reward or accuracy rises while traces become less grounded and monitorable. Prior work largely documents this decay behaviorally, leaving its representation-level correlates and actionable controls unclear. We find that template- and ground-associated activations become less separable during RL; matched interventions support the contribution of selected features to monitorability degradation. Guided by this evidence, we propose Targeted Anti-obfuscation with Mechanistic Enforcement (TAME), which uses Sparse Autoencoders (SAEs) to combine behavioral feedback with targeted suppression of template-associated activations during RL. Its asymmetric constraint penalizes template activations only above their pre-RL baseline, anchoring the localized features while behavioral feedback promotes grounded refinements. Across VIRL-39k, SPA-VL, and two model families, TAME improves CoT monitorability by up to 30.9 and 16.7 percentage points over Group Relative Policy Optimization (GRPO), respectively. Blinded human evaluation finds higher human monitorability on both datasets, and two held-out monitor families reproduce the monitorability gains. Task accuracy changes are small and mixed, and general-capability benchmarks show task-specific trade-offs. These results provide a path from behavioral monitoring to representation-level oversight for more auditable RL-trained multimodal systems.