arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LEMUR:基于视觉锚定推理重定向的潜在熵感知多模态遗忘

LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection

Xinhao Zhong, Yuxia Qiao, Junhao Li, Hao Fang, Yi Sun, Bin Chen

arXiv 2608.11691首次发表:更新:

AI 中文总结

针对原生RL训练多模态模型的隐私泄漏问题,提出基于熵动态的无训练推理时遗忘框架LEMUR,可更有效地抑制敏感信息在推理轨迹和答案中的泄漏,同时保留模型非敏感效用。

AI 中文摘要

强化学习(RL)后训练为多模态大推理模型(MLRMs)赋予了探索性思维链(CoT),大幅提升了视觉推理能力。然而,我们发现该能力引入了一种独特的隐私漏洞:即使从最终答案中成功遗忘了敏感事实,模型仍可能在其推理轨迹中重现该事实。这种泄漏在原生RL训练的MLRMs中比在无推理能力的基础模型中显著得多,揭示了现有遗忘方法未设计解决的隐私风险。我们表明,RL诱导的探索使敏感内容具有独特的token级熵特征,而基础模型中基本不存在该特征。基于此观察,我们提出LEMUR,一种针对原生RL训练多模态模型的完全无训练、推理时的遗忘框架。LEMUR使用熵动态作为控制信号,识别敏感推理何时开始以及何时应停止清理。在此区间内,它通过熵调制的视觉锚定潜在注入重定向推理轨迹,用重新基于输入图像的、经概率加权的清理嵌入替换已确定的token。在多种MLRMs上,LEMUR在抑制推理轨迹和答案泄漏方面始终优于现有遗忘方法,同时更好地保留非敏感效用和输出流畅性。这些结果表明,RL诱导的熵动态为隐私泄漏提供了独特信号,利用该信号可实现针对具备推理能力的多模态模型的有效无训练遗忘。

英文摘要

Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, we find that this capability introduces a distinct privacy vulnerability: even when a sensitive fact is successfully unlearned from the final answer, the model may still reproduce it in its reasoning trace. This leakage is substantially more pronounced in natively RL-trained MLRMs than in their non -reasoning base models, revealing a privacy risk that existing unlearning methods are not designed to address. We show that RL-induced exploration leaves sensitive content with a distinctive token-level entropy signature that is largely absent from base models. Based on this observation, we propose LEMUR, a fully training-free, inference-time unlearning framework for natively RL-trained multimodal models. LEMUR uses entropy dynamics as a control signal to identify when sensitive reasoning begins and when sanitization should stop. During this interval, it redirects the reasoning trajectory through entropy-modulated visual-anchor latent injection, replacing committed tokens with sanitized, probability-weighted embeddings re-grounded in the input image. Across diverse MLRMs, LEMUR consistently outperforms existing unlearning met hods in suppressing both reasoning-trace and answer leakage, while better preserving non-sensitive utility and output fluency. These results demonstrate that RL-induced entropy dynamics provide a distinctive signal for privacy leakage and that exploiting this signal enables effective training-free unlearning for reasoning-capable multimodal models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑