评估潜在推理模型的对抗鲁棒性
Assessing Adversarial Robustness of Latent Reasoning Models
查看机构详情
- Yuanpei College, Peking University(北京大学元培学院)
- CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究系统评估了潜在推理模型在文本和多模态场景下的对抗鲁棒性,发现其普遍不如显式CoT基线,尤其在白盒攻击下性能严重下降,并揭示了不同模态的失败模式,强调设计隐式推理系统需兼顾效率与鲁棒性。
中文摘要 AI 辅助
大型语言模型越来越依赖长链式思维(CoT)轨迹进行复杂推理,但自回归生成带来了大量的内存和推理成本。潜在推理模型(LRMs)通过将中间推理压缩为少量连续的潜在向量,提供了一种更高效的替代方案。尽管效率高,但LRMs的对抗鲁棒性在很大程度上仍未得到充分探索。在这项工作中,我们系统地评估了潜在推理在文本和多模态设置中的鲁棒性,涵盖了八个模型和六个基准。我们发现,在我们评估的设置中,LRMs在对抗扰动下通常不如显式CoT基线鲁棒,尤其是在白盒攻击下性能严重下降。进一步的分析揭示了不同模态的失败模式:文本潜在状态表现出脆弱的动态性,对特定输入模式高度敏感,而多模态模型中的潜在状态可能对输入扰动基本保持不变,对最终预测的影响有限。这些发现暴露了当前潜在推理方法的鲁棒性局限,并强调了在设计隐式推理系统时需要同时考虑效率和鲁棒性。我们已开源代码以促进本研究的复现,链接为https://this URL。
英文摘要
Large language models increasingly rely on long chain-of-thought (CoT) trajectories for complex reasoning, but autoregressive generation brings substantial memory and inference costs. Latent reasoning models (LRMs) offer a more efficient alternative by compressing intermediate reasoning into a small number of continuous latent vectors. Despite their efficiency, however, the adversarial robustness of LRMs remains largely underexplored. In this work, we systematically evaluate the robustness of latent reasoning across textual and multimodal settings, covering eight models and six benchmarks. We find that, across our evaluated settings, LRMs are generally less robust than explicit CoT baselines under adversarial perturbations, with particularly severe degradation under white-box attacks. Further analysis reveals distinct failure modes across modalities: textual latent states exhibit brittle dynamics and high sensitivity to specific input patterns, while latent states in multimodal models can remain largely invariant to input perturbations and have limited influence on final predictions. These findings expose robustness limitations of current latent reasoning approaches and highlight the need to jointly consider efficiency and robustness when designing implicit reasoning systems. We have open-sourced our code to facilitate reproduction of our research https://github.com/PKU-ML/latent-reasoning-model-assessment.