arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于选择的结构化推理:面向高效多模态搜索智能体

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang

arXiv 2610.01892首次发表:更新:

发表机构

Apple; Carnegie Mellon University(苹果公司; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于选择的结构化推理(SSR),将多模态智能体的推理从开放式生成改为选择预定义候选,在保持任务性能的同时显著降低推理延迟,适用于小型模型。

AI 中文摘要

多模态智能体通常在每次行动前生成自由形式的推理。对于小型模型,有限的模型容量可能导致冗长的推理,这些推理对行动生成提供的指导甚少,同时产生大量推理成本。为应对这一挑战,我们引入了基于选择的结构化推理(SSR),一种将推理重新定义为选择而非开放式生成的框架。SSR将重复出现的高层推理表示为预先指定的、可复用的自然语言候选。在每一轮中,模型根据当前上下文下的似然性从这些推理候选中进行选择,无需辅助任务头。使用预先指定的推理轨迹支持并行评分,其中教师强制预填充利用共享上下文KV缓存,在候选内部及候选之间并发计算令牌似然。我们在七个多模态搜索基准上使用2B和4B模型评估了SSR。在多种强化学习目标和监督微调下,SSR在不牺牲任务性能的情况下实现了显著的效率提升。SSR的平均成功率与同规模领先的搜索智能体相当,同时将每轮推理延迟降低超过90%,每问题总模型推理延迟降低28-54%。项目页面:此https URL。

英文摘要

Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: https://zfy0314.github.io/ssr-webpage/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑