arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12781cs.CV

超越正确性:混合思维多模态大语言模型(MLLM)的响应行为基准测试与对齐

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

Xinming Wang, Weinong Wang, Hongming Yang, Yansong Lin, Zheng Ruan, Shangpin Peng, Qiming Peng, Nan Qiao, Fengyuan Lu, Guoqing Ma, Marito Li, Songyang Zhang, Sa… 展开作者

Xinming Wang, Weinong Wang, Hongming Yang, Yansong Lin, Zheng Ruan, Shangpin Peng, Qiming Peng, Nan Qiao, Fengyuan Lu, Guoqing Ma, Marito Li, Songyang Zhang, Saiyong Yang, Han Hu, Yonglong Tian, Xu-Yao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对混合思维 MLLM 的思维与非思维模式响应错位问题,构建 PatternEval 基准并开发 PatternRL 方法,可减轻跨模式错位且任务性能损失极小。

中文摘要 AI 辅助

混合思维多模态大语言模型(MLLM)允许单个模型在 deliberative 思维模式与低延迟的非思维推理模式之间切换。尽管这些模式的推理预算不同,但其输出的响应应满足相同的用户面向标准。仅正确性不足以表征该响应质量;因此,我们将任务准确率和响应模式故障作为互补结果进行评估。我们通过响应模式对齐研究这一差距:即思维和非思维接口是否保持可接受的最终响应行为。我们推出 PatternEval,这是一个包含2415个多模态提示的故障增强诊断基准,涵盖视觉感知与 grounding、结构化图像理解以及多模态知识推理三大类任务。PatternEval 测试四种反复出现的故障:思维链泄露、响应重复、逻辑矛盾和表演性推理。响应模式故障在不同提供商的模型中普遍存在,其中非思维推理的故障率显著更高,从而在思维与非思维接口之间造成系统性错位。基于这一诊断,我们开发 PatternRM,一种响应级奖励模型,以及 PatternRL,其在强化学习期间引入模式特定惩罚。在 Qwen3-VL-4B 和 Qwen3-VL-8B 上的实验表明,在强化学习中纳入模式特定惩罚可减轻跨模式错位,同时仅产生微小的任务性能权衡。总体而言,PatternEval 和 PatternRL 为跨混合思维接口对齐用户可见的响应模式提供了评估与训练框架。

英文摘要

Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not characterize this response quality; we therefore evaluate task accuracy and response-pattern failures as complementary outcomes. We study this gap through \textbf{response-pattern alignment}: whether thinking and non-thinking interfaces preserve acceptable final-response behavior. We introduce \textbf{PatternEval}, a failure-enriched diagnostic benchmark comprising 2,415 multimodal prompts spanning visual perception and grounding, structured image understanding, and multimodal knowledge reasoning. PatternEval tests four recurrent failures: chain-of-thought leakage, response repetition, logical contradiction, and performative reasoning. Response-pattern failures are widespread across models from different providers, with non-thinking inference exhibiting substantially higher failure rates and thereby creating systematic misalignment between thinking and non-thinking interfaces. Motivated by this diagnosis, we develop \textbf{PatternRM}, a response-level reward model, and \textbf{PatternRL}, which introduces pattern-specific penalties during reinforcement learning. Experiments on Qwen3-VL-4B and Qwen3-VL-8B show that incorporating pattern-specific penalties into reinforcement learning can mitigate cross-mode misalignment while incurring a marginal task performance trade-off. Together, PatternEval and PatternRL provide an evaluation-and-training framework for aligning user-visible response patterns across hybrid-thinking interfaces.

发表机构

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • Large Language Model Department, Tencent(腾讯大语言模型部)
  • University of Electronic Science and Technology of China(电子科技大学)
  • Hong Kong University of Science and Technology(香港科技大学)
  • Zhongguancun Academy(中关村学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑