arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体视觉语言流水线中的视觉编排税:视觉证据重用的审计与认证

Visual Orchestration Tax in Agentic VLM Pipelines: Auditing and Certifying Visual Evidence Reuse

Lingteng Zeng

arXiv 2610.08170首次发表:更新:

发表机构

The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体VLM流水线中静态视觉证据重复传递造成的编排冗余,提出审计与认证框架,量化冗余并引入SharedVisCache实现行为保持的视觉重用,显著降低视觉计算开销。

AI 中文摘要

智能体视觉语言(VLM)流水线日益将相同的静态视觉证据传递给多个专业智能体和工具。这种设计产生了一种编排层面的冗余模式:语义未改变的图像在VLM API边界被反复重建为图像条件请求。我们将这一现象称为视觉编排税,并开发了一个面向智能体VLM流水线中视觉证据重用的从测量到认证的框架。审计方面定义了M1_trace来统计原始视觉证据接触次数,以及M2来度量结构性接触冗余,并配有查询级分布、自助置信区间和配对质量检验。在SeeingEye和MAMMQA上针对图表、文档、通用视觉问答和多模态问答任务的审计显示,视觉证据接触冗余度为66.8%-75.6%,且每个被审计的查询均超过预设门槛。认证方面引入了SharedVisCache,这是一个基于图像内容、预处理指纹和编码器假设的契约感知证据重用钩子。在SeeingEye上,契约验证认证75.0%-75.5%的重复接触为可重用,同时保持350/350的输出字符串和ΔM5=0。在物理层面,认证命中将ChartQA-200轨迹重放中的F_vision从800降至200,并在实时SeeingEye翻译器阶段物理集成中从200降至50,同时保持800/800的重放字符串和200/200的集成调用输出。这些结果将视觉重用定位为智能体编排的可测量、行为保持属性,并定义了一个智能体层契约,使后端前缀或令牌重用具有语义可解释性。

英文摘要

Agentic VLM pipelines increasingly pass the same static visual evidence through multiple specialist agents and tools. This design creates an orchestration-level redundancy mode: semantically unchanged images are repeatedly reconstructed as image-conditioned requests at the VLM API boundary. We call this phenomenon visual orchestration tax and develop a measurement-to-certification framework for visual evidence reuse in agentic VLM pipelines. The audit side defines $\mathrm{M1}_{\mathrm{trace}}$ to count raw visual-evidence touches and M2 to measure structural touch redundancy, with query-level distributions, bootstrap confidence intervals, and paired quality tests. Across SeeingEye and MAMMQA on chart, document, general-VQA, and multi-modal-QA tasks, audits reveal 66.8-75.6% visual-evidence touch redundancy, and every audited query exceeds the predefined gate. The certification side introduces SharedVisCache, a contract-aware evidence reuse hook keyed by image content, preprocessing fingerprint, and encoder assumptions. On SeeingEye, contract validation certifies 75.0-75.5% repeated touches as reusable while preserving 350/350 output strings and $Δ\mathrm{M5}{=}0$. At the physical layer, certified hits reduce $F_{\mathrm{vision}}$ from 800 to 200 in ChartQA-200 trace replay and from 200 to 50 inside live SeeingEye translator-stage physical integration, preserving 800/800 replay strings and 200/200 integrated call outputs. The results position visual reuse as a measurable, behavior-preserving property of agent orchestration and define an agent-layer contract that makes backend prefix or token reuse semantically interpretable.

Comments9 pages, 2 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑