发表机构
NLPR&MAIS, Institute of Automation, Chinese Academy of Sciences; School of Advanced Interdisciplinary Science, University of Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Zhongguancun Academy(中国科学院自动化研究所模式识别国家重点实验室与多模态人工智能系统实验室; 中国科学院大学前沿交叉科学学院; 中国科学院大学人工智能学院; 中关村学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对沙盒环境中VLM视觉推理的状态管理难题,提出无训练框架VLM-in-Sandbox,通过视觉工作区分离证据生成与管理,在七个基准上取得最优准确率并降低token开销。
AI 中文摘要
沙盒化计算机环境支持使用工具、可执行程序和持久化文件进行多步推理,然而将其从语言模型扩展到视觉语言模型(VLM)引入了独特的状态管理问题。视觉推理会产生中间图像型证据——裁剪、掩码、叠加、缩放区域和分析渲染图——这些证据必须保持可寻址,同时避免在多模态上下文中无界累积。我们提出VLM-in-Sandbox,一种在受控计算机环境中进行智能体多模态推理的无训练框架。其视觉工作区在图像账本中注册生成的工件,维护有界的活动视觉上下文,并允许模型显式提升选定的证据以供后续检查。这将视觉证据生成(由沙盒工具执行)与视觉证据管理分离开来。在七个基准和四个基础VLM上,VLM-in-Sandbox在Vanilla VLM、Append-only Sandbox和所提方法中取得了最高的样本加权平均准确率。一项编译器匹配的2×2研究在1,260个示例上进一步区分了模型引导的可见性与有界保留:VLM-in-Sandbox达到66.27%的准确率,同时总token数比自动保留全部的对照方法少18.6%。在所有6,350个提交的GPT-4.1-mini示例中,相对于原始Append-only,它产生了302次拯救和142次回退。一项使用前缀缓存的本地vLLM研究证实,较小的请求工作负载也减少了未缓存token、首token时间和端到端延迟。这些结果将显式视觉证据状态确定为沙盒化VLM智能体的核心抽象。
英文摘要
Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language models to vision-language models (VLMs) introduces a distinct state-management problem. Visual reasoning produces intermediate image-valued evidence---crops, masks, overlays, zoomed regions, and analytic renderings---that must remain addressable without accumulating unboundedly in multimodal context. We introduce VLM-in-Sandbox, a training-free framework for agentic multimodal reasoning in controlled computer environments. Its Visual Workspace registers generated artifacts in an image ledger, maintains a bounded active visual context, and lets the model explicitly promote selected evidence for subsequent inspection. This separates visual evidence generation, performed by sandbox tools, from visual evidence management. Across seven benchmarks and four base VLMs, VLM-in-Sandbox achieves the highest sample-weighted average accuracy among Vanilla VLM, Append-only Sandbox, and the proposed method. A compiler-matched $2\times2$ study on 1,260 examples further separates model-directed visibility from bounded retention: VLM-in-Sandbox reaches 66.27% accuracy with 18.6% fewer total tokens than the automatic, retain-all control. Over all 6,350 submitted GPT-4.1-mini examples, it produces 302 rescues and 142 regressions relative to Original Append-only. A local vLLM study with prefix caching confirms that the smaller request workload also reduces uncached tokens, time to first token, and end-to-end latency. These results identify explicit visual evidence state as a central abstraction for sandboxed VLM agents.