arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

回答我的问题需要多少成本?基于云VLM的VQA系统基准测试

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems

Henri Vanhuynegem, Weitao Xu, Yiran Shen, Guohao Lan

arXiv 2608.07861首次发表:更新:

AI 中文总结

本研究推出首个将客户端输入预处理作为受控变量的VQABench基准,评估12种预处理技术在3个VQA数据集、4个商业VLMs上的95168次API调用,明确预处理对云VLM-based VQA的成本-质量影响,为VQA系统部署提供指导。

AI 中文摘要

视觉语言模型(VLMs)正成为移动视觉问答(VQA)系统的实用后端,使智能手机和智能眼镜能够回答用户关于物理世界的问题。由于现代VLMs仍难以在移动和边缘设备上运行,VQA系统越来越多地将推理任务卸载到基于云的VLMs上。这让移动设备获得了更强的计算能力,但也使得视觉输入准备成为关键的系统变量:卸载前图像的准备方式不仅影响答案质量,还影响有效载荷大小、令牌成本和系统延迟。专有API对模型内部或服务行为几乎没有控制权,使得客户端预处理成为下游开发者主要的实际优化空间。许多视觉卸载技术已被提出,但它们对商业云VLMs的成本-质量影响从未被研究。为填补这一空白,我们推出VQABench,这是首个将客户端输入预处理作为云VLM-based VQA的受控变量的系统基准。我们在三个VQA数据集和三个提供商的四个商业VLMs上评估了12种预处理技术,共进行95,168次API调用。结果表明,预处理并非普遍有益:其有效性取决于目标模型、API范式、提供商的令牌核算规则和任务形式。选择不当的预处理策略会增加部署成本或延迟,同时降低答案准确性。总体而言,我们的基准阐明了预处理何时有效、何时失效及原因,为VQA系统的未来研究和实际部署提供了见解。

英文摘要

Vision-language models (VLMs) are becoming a practical backend for mobile visual question answering (VQA) systems, enabling smartphones and smart glasses to answer users' questions about the physical world. Since modern VLMs remain difficult to run on mobile and edge devices, VQA systems increasingly offload inference to cloud-based VLMs. This gives mobile devices access to stronger computation, but it also makes visual input preparation a key system variable: how the image is prepared before offloading affects not only answer quality but also payload size, token cost, and system latency. Proprietary APIs expose little control over model internals or serving behavior, leaving client-side preprocessing as the main practical optimization space for downstream developers. Many such techniques have been proposed for visual offloading, yet their cost-quality impact on commercial cloud VLMs has never been studied. To fill this gap, we present VQABench, the first systematic benchmark that treats client-side input preprocessing as a controlled variable for cloud-VLM-based VQA. We evaluate 12 preprocessing techniques across three VQA datasets and four commercial VLMs from three providers, totaling 95,168 API calls. Our results show that preprocessing is not universally beneficial: its effectiveness depends on the target model, API paradigm, provider token-accounting rule, and task formulation. A poorly selected preprocessing strategy can increase deployment cost or latency while degrading answer accuracy. Overall, our benchmark clarifies when preprocessing helps, when it fails, and why, providing insights to guide future research and real-world deployment of VQA systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑