arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09313cs.CV

为什么视觉语言模型会漏检小目标,以及何时放大是安全的

Why VLMs Miss Small Objects, and When Zooming In Is Safe

  • University of California, Berkeley(加州大学伯克利分校)
  • Quotr AI

机构由 AI 辅助整理,请以论文原文为准。

Junzhe Shi, Yuan Gan, Shida Jiang

AI总结:

本研究构建理论解释VLM漏检小目标的原因,证明token预算与覆盖成本的权衡,并验证局部分解在特定条件下不降低召回率,且能提升性能。

AI中文摘要:

视觉语言模型(VLMs)在处理大图像时经常漏检小目标。我们提出三个问题:是什么限制了它们,其中哪些限制是更好的模型可以消除的,以及处理大图像的经典方法——局部分解——是否仍有未来。我们基于图像接口的两个量构建了一个理论来回答这些问题:S,即一个目标边上视觉 token 的数量,以及 L,即一次调用必须覆盖的内容。限制如下:一张 W x H 的图像在 N 个 token 内整体送入时,一个边长为 m 的目标每边最多获得 m*sqrt(N/(WH)) 个 token。因此,将 token 预算 N 加倍,整体图像能达到的最大 S 仅提高 41%,而无论模型如何,以目标所需的 S 查看图像至少需要 S^2 量级的 token。如果识别性能随 S 逐步提升,那么任何搜索策略(包括缩放智能体)都遵循一个召回-成本边界。更好的模型能改变的是:目标所需的 S 以及一次调用能携带的内容(以每个找到的目标的比特数衡量);覆盖成本保持不变。分解:是的。仅假设每个目标更多 token 和每次调用更少内容平均不会有害,那么如果没有任何视图相对于整体图像缩小,且视图间重叠一个目标,则分割图像不会降低召回率。该条件的任一部分都不能省略;我们限定了每种此类分解的成本,并且随着图像增大,一个简单规则接近该边界。我们在 797 张图像上进行了约 177,000 次请求来测试该理论。在受控图像上,为 8 个 VLM 预测的 40 个排序中没有一个被违反。事后在图纸、平面图、自然图像和合成图像上检查,89 个隐含排序中有 61 个显著成立,4 个失败,且失败均出现在 OpenAI 模型获得比默认路径更多像素的情况下。在施工图纸上,该规则相对于整体图像从未显著降低召回率,并最多提高了 0.28。代码和数据:此 https URL

英文摘要:

Vision-language models (VLMs) often miss small objects in large images. We ask three questions: what limits them, which of these limits better models can remove, and whether the classical way of handling large images, local decomposition, still has a future. We answer them with a theory built on two quantities of the image interface: S, the number of visual tokens across an object's side, and L, the content a call must cover. The limits: a W x H image sent whole within N tokens gives an object of side m at most m*sqrt(N/(WH)) tokens per side. Doubling the token budget N therefore raises the largest S a whole image can reach by only 41%, and seeing an image at the S an object needs costs at least order S^2 tokens whatever the model. If recognition improves gradually with S, any search strategy, zoom agents included, obeys a recall-cost frontier. What better models can change: the S an object needs and how much one call can carry, measured in bits per object found; the coverage cost remains. Decomposition: yes. Assuming only that more tokens per object and less content per call do not hurt on average, splitting an image cannot lower recall if no view zooms out relative to the whole image and views overlap by one object. Neither part of this condition can be dropped; we bound the cost of every such decomposition, and a simple rule approaches the bound as the image grows. We test the theory in about 177,000 requests on 797 images. On controlled images, none of 40 orderings predicted for 8 VLMs is violated. Checked after the fact on drawings, floor plans, natural and synthetic images, 61 of 89 implied orderings hold significantly and 4 fail, all for OpenAI models given more pixels than their default path. On construction drawings the rule never significantly lowered recall relative to the whole image and raised it by up to 0.28. Code and data: https://github.com/shijunzhe/vlm-small-objects

↑