超越正确性:混合思维多模态大语言模型(MLLM)的响应行为基准测试与对齐
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
浏览论文内容
中文总结 AI 辅助
该研究针对混合思维 MLLM 的思维与非思维模式响应错位问题,构建 PatternEval 基准并开发 PatternRL 方法,可减轻跨模式错位且任务性能损失极小。
中文摘要 AI 辅助
混合思维多模态大语言模型(MLLM)允许单个模型在 deliberative 思维模式与低延迟的非思维推理模式之间切换。尽管这些模式的推理预算不同,但其输出的响应应满足相同的用户面向标准。仅正确性不足以表征该响应质量;因此,我们将任务准确率和响应模式故障作为互补结果进行评估。我们通过响应模式对齐研究这一差距:即思维和非思维接口是否保持可接受的最终响应行为。我们推出 PatternEval,这是一个包含2415个多模态提示的故障增强诊断基准,涵盖视觉感知与 grounding、结构化图像理解以及多模态知识推理三大类任务。PatternEval 测试四种反复出现的故障:思维链泄露、响应重复、逻辑矛盾和表演性推理。响应模式故障在不同提供商的模型中普遍存在,其中非思维推理的故障率显著更高,从而在思维与非思维接口之间造成系统性错位。基于这一诊断,我们开发 PatternRM,一种响应级奖励模型,以及 PatternRL,其在强化学习期间引入模式特定惩罚。在 Qwen3-VL-4B 和 Qwen3-VL-8B 上的实验表明,在强化学习中纳入模式特定惩罚可减轻跨模式错位,同时仅产生微小的任务性能权衡。总体而言,PatternEval 和 PatternRL 为跨混合思维接口对齐用户可见的响应模式提供了评估与训练框架。
英文摘要
Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not characterize this response quality; we therefore evaluate task accuracy and response-pattern failures as complementary outcomes. We study this gap through \textbf{response-pattern alignment}: whether thinking and non-thinking interfaces preserve acceptable final-response behavior. We introduce \textbf{PatternEval}, a failure-enriched diagnostic benchmark comprising 2,415 multimodal prompts spanning visual perception and grounding, structured image understanding, and multimodal knowledge reasoning. PatternEval tests four recurrent failures: chain-of-thought leakage, response repetition, logical contradiction, and performative reasoning. Response-pattern failures are widespread across models from different providers, with non-thinking inference exhibiting substantially higher failure rates and thereby creating systematic misalignment between thinking and non-thinking interfaces. Motivated by this diagnosis, we develop \textbf{PatternRM}, a response-level reward model, and \textbf{PatternRL}, which introduces pattern-specific penalties during reinforcement learning. Experiments on Qwen3-VL-4B and Qwen3-VL-8B show that incorporating pattern-specific penalties into reinforcement learning can mitigate cross-mode misalignment while incurring a marginal task performance trade-off. Together, PatternEval and PatternRL provide an evaluation-and-training framework for aligning user-visible response patterns across hybrid-thinking interfaces.
发表机构
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- Large Language Model Department, Tencent(腾讯大语言模型部)
- University of Electronic Science and Technology of China(电子科技大学)
- Hong Kong University of Science and Technology(香港科技大学)
- Zhongguancun Academy(中关村学院)
机构由 AI 辅助整理,请以论文原文为准。