arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

思想令牌混合:统一感知与推理以实现自由形式多模态基础定位

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

Tianyi Gao, Han Fang, Tianyi Ding, Hao Li, Xin Wei, Hongbo Sun, Xiaodong Dong, Ye Yuan, Jinglin Xu, Kongming Liang, Hao Sun, Jingmin Xin

arXiv 2607.24407首次发表:更新:

发表机构

Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University; Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd; University of Science and Technology Beijing; Beijing University of Posts and Telecommunications; Shanghai Jiao Tong University(西安交通大学人工智能与机器人研究所; 星辰通用人工智能实验室,中国电信人工智能技术(北京)有限公司; 北京科技大学; 北京邮电大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在统一多模态基础定位中的感知与推理。提出思想令牌混合(Motto)方法,通过空间接地思想令牌化及上下文自适应令牌链实现,构建PR-Bench基准,实验证明该方法在自由形式接地任务中达领先性能。

AI 中文摘要

多模态大语言模型在基础定位任务中取得了很大进展,但现有方法仍难以统一精确的定位和复杂的推理。基于文本的方法依赖坐标或索引预测,限制了模型对密集视觉对象的感知能力;基于潜在令牌的方法使用缺乏固有空间参考的特殊令牌且解码机制缺乏思维步骤,削弱了高级推理能力。为此,我们提出了思想令牌混合(Motto)方法,它通过空间接地思想令牌化使特殊令牌与空间位置明确对齐,设计上下文自适应令牌链在交错推理链中动态切换接地模式,还构建了PR-Bench基准。实验表明Motto在各种自由形式接地任务中实现了领先性能。

英文摘要

Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing, text-based methods rely on coordinates or index prediction, severely limiting the perceptual capabilities of the model for dense visual objects. Meanwhile, latent token-based methods employ special tokens without inherent spatial references and use a decoding mechanism that lacks thinking steps, weakening high-level reasoning capabilities. Consequently, developing a unified framework that excels in both perception and reasoning remains challenging. To address this, we propose Mixture-of-Thought-Tokens (Motto), a new free-form multimodal grounding method that bridges the perception-reasoning gap, enabling MLLMs to empower diverse, arbitrary grounding queries. Specifically, we introduce Spatially-Grounded Thought Tokenization to explicitly align special tokens with spatial locations for clear spatial correspondence and visual interpretability. We further design a Context-Adaptive Chain-of-Tokens that dynamically switch grounding modes within an interleaved reasoning chain, achieving robust grounding across tasks of varying complexity. In addition, we construct PR-Bench, a new referring expression comprehension benchmark to evaluate the perception-reasoning gap. Extensive experiments demonstrate that Motto achieves state-of-the-art performance across diverse free-form grounding tasks.

CommentsAccepted by ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑