发表机构
Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大语言模型智能体长轨迹中文本主导上下文、抑制视觉证据的问题,提出基于KL引导的SPARE框架,结合OPSD与SFT实现高效剪枝,在多步视觉工具使用基准中达到最优准确率与剪枝比例的权衡。
AI 中文摘要
多模态大语言模型(MLLM)正越来越多地被部署为多步智能体,其中显式推理支持任务分解与工具协调,但也会积累自身生成的文本。在长轨迹中,这些文本会主导上下文并抑制视觉证据,形成文本债务。我们观察到,一旦与任务相关的视觉证据被定位,推理就会变得冗余;而当定位仍不确定时,过时的假设会误导后续推理。因此,剪枝必须移除冗余文本而不丢弃视觉证据。我们提出SPARE,一种基于Kullback–Leibler(KL)的框架,用于剪枝多模态工具使用智能体中积累的推理内容。SPARE使用紧凑的任务状态摘要作为特权诊断上下文。对于每个候选片段,它在原始上下文和摘要条件化上下文下重放同一模型。来自在线策略自蒸馏(OPSD)的反向KL散度随后测试摘要是否充分覆盖该片段而不破坏未来推理。我们进一步通过监督微调(SFT)微调摘要器,使其能生成更紧凑、覆盖范围更广的摘要,并进行更激进的剪枝。在多步视觉工具使用基准测试中,SPARE在剪枝方法中实现了最高的平均准确率,同时移除了37.89%至64.58%的推理标记。这种有利的准确率-上下文权衡表明,减少文本主导地位可恢复对视觉证据的依赖,并减轻对自身生成语言的过度条件化。
英文摘要
Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit reasoning supports task decomposition and tool coordination but also accumulates self-generated text. Over long trajectories, this text can dominate the context and suppress visual evidence, creating textual debt. We observe that reasoning becomes redundant once task-relevant visual evidence is grounded, while stale hypotheses can misguide later inference when grounding remains uncertain. Pruning must therefore remove redundant text without discarding visual evidence. We propose SPARE, a Kullback-Leibler (KL)-guided framework for pruning accumulated reasoning in multimodal tool-use agents. SPARE uses a compact task-state summary as privileged diagnostic context. For each candidate segment, it replays the same model under the original and summary-conditioned contexts. Reverse-KL divergence from on-policy self-distillation (OPSD) then tests whether the summary sufficiently covers the segment without disrupting future reasoning. We further fine-tune the summarizer with supervised fine-tuning (SFT), enabling more compact summaries, broader coverage, and more aggressive pruning. Across multi-step visual tool-use benchmarks, SPARE achieves the highest average accuracy among pruning methods while removing 37.89-64.58\% of reasoning tokens. This favorable accuracy-context trade-off shows that reducing textual dominance restores reliance on visual evidence and mitigates over-conditioning on self-generated language.
Comments17 pages, 3 figures, 5 tables