发表机构
Hong Kong University of Science and Technology; The Chinese University of Hong Kong; Ant Group; Zhejiang University(香港科技大学; 香港中文大学; 蚂蚁集团; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出工具调用支架假设,引入TextCall方法验证返回图像冗余,仅用结构化文本支架即可实现视觉推理,降低延迟并消除工具API调用,结论适用于现有相关基准。
AI 中文摘要
工具增强型视觉语言模型越来越多地“用图像思考”:它们调用裁剪、缩放或代码工具,并对返回的像素进行推理。然而,近期采用盲测、增益分解和注意力分析的研究表明,返回的图像贡献甚微,这就提出了一个问题:如果像素没有带来增益,那什么才是关键?我们假设,承载关键信息的信号是在任何返回像素到达之前发出的结构化文本:工具名称、坐标、目标描述和意图。这种文本支架编码了应查看的位置和要查找的内容。我们引入TextCall(仅调用不返回)来验证这一假设:它保留支架,但将返回的图像替换为文本占位符“[Image output skipped]”。三项研究支持该假设。(i)返回像素的非必要性:在LoRA、全微调、RL三种设置下,TextCall的表现与完整的“用图像思考”相当甚至更好;在RL设置下,它在报告的检查点保留了工具使用能力,避免了在匹配设置下,看到返回图像会导致模型停止调用工具并直接回答的失效模式。(ii)支架的充分性:在匹配的训练查询上,仅支架输入的准确率与返回图像输入的准确率相当。(iii)组件特异性:将支架分解为推理文本和空间代码后,两个组件均有贡献,且主导组件因任务而异。这些结果共同支持工具调用支架假设:在当前“用图像思考”的分布中,活跃信号是工具调用时发出的结构化文本;返回的图像是冗余载体。TextCall在保持准确率的同时,将延迟降低了29%-46%,并消除了工具执行API调用。我们的结论适用于当前的“用图像思考”基准;构建像素真正承载关键信息的任务仍是一个开放方向。
英文摘要
Tool-augmented vision-language models increasingly "think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind tests, gain decompositions, and attention analyses has shown that returned images contribute little, raising the question: if pixels do not carry the gain, what does? We hypothesize that the load-bearing signal is the structured text emitted before any returned pixel arrives: tool name, coordinates, target description, and intent. This textual scaffold encodes where to look and what to find. We introduce TextCall (call-but-no-return) to test this: it keeps the scaffold but replaces returned images with the text placeholder [Image output skipped]. Three studies support the hypothesis. (i) Non-necessity of returned pixels: across LoRA, full fine-tuning, and RL, TextCall matches or exceeds full thinking-with-images; under RL it preserves tool use at the reported checkpoint, avoiding the failure mode where, under matched settings, seeing the returned image causes the model to stop calling tools and answer directly. (ii) Sufficiency of the scaffold: on matched training queries, scaffold-only input yields equivalent accuracy to returned-image input. (iii) Component specificity: decomposing the scaffold into reasoning text and spatial code shows both components contribute, with the dominant one varying by task. Together these results support the Tool-Call Scaffold Hypothesis: in current thinking-with-images distributions, the active signal is the structured text emitted at tool-call time; the returned image is a redundant carrier. TextCall preserves accuracy while reducing latency by 29-46% and eliminating tool-execution API calls. Our claims hold for current thinking-with-images benchmarks; constructing tasks where pixels are genuinely load-bearing remains an open direction.