arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32757cs.AIcs.CVcs.LG

读出并非恢复:在视觉-语言模型中解离坐标发射与视觉损坏修复

Readout is not Recovery: Dissociating Coordinate Emission from Visual-Corruption Repair in Vision-Language Models

Drandreb Earl Juanico

AI总结:

本研究通过Qwen3-VL和Kimi-VL实验,发现坐标令牌读出与视觉损坏修复在层间解离,端点坐标分流可作为电路先验,但遮挡恢复需独立基准。

AI中文摘要:

VLM边界框定位既是语言生成也是空间承诺。诸如bbox_2d之类的可解析字段使定位易于评分,但发射坐标令牌的维度在视觉证据受损后无需修复定位。我们在Qwen3-VL-4B-Instruct上针对单目标COCO接地研究这种读出/恢复分离。我们比较了干净坐标令牌读出排名与源自损坏的修复排名,使用对象掩码端点替换进行恢复,并使用干净输入下限进行深度定位。在Qwen3-VL中,坐标令牌排名在第24层之前保持惰性,在第32-35层承担负载,并在第34层达到峰值;源自损坏的排名在第16-24层有害,但在第35层/最终层附近变得有益。Kimi-VL-A3B诊断显示,尽管框格式不同,但存在匹配的输出近端转换。对象掩码恢复分离了排名预算:$k=250$显示必要性,$k=500$显示高于随机的Top-$k$恢复,而$k=d/2$主要受容量驱动。部分遮挡扫描揭示,高重叠坐标令牌集在$k=1000$时可能有害,主要在半宽度时有益,而群体源自损坏的集不提供可靠的固定修复集。边缘归因修补显示坐标令牌路径对检测恢复具有高精度但低召回率,而RMSNorm准层控制并未缩小端点修复差距。因此,端点坐标分流是有用的电路先验,但遮挡恢复需要单独的基准。

英文摘要:

VLM bounding-box localization is both language generation and spatial commitment. Parseable fields such as bbox_2d make localization easy to score, but dimensions that emit coordinate tokens need not repair localization after visual evidence is damaged. We study this readout/recovery separation in Qwen3-VL-4B-Instruct on single-object COCO grounding. We compare clean coordinate-token readout rankings with corruption-derived repair rankings, using object-mask endpoint replacement for recovery and clean-input flooring for depth localization. In Qwen3-VL, coordinate-token rankings are inert through layer 24, load-bearing from layers 32-35, and peak at layer 34; corruption-derived rankings harm layers 16-24 but become beneficial near layer 35/final. A Kimi-VL-A3B diagnostic shows a matching output-proximal transition despite a different box format. Object-mask recovery separates rank budgets: $k=250$ shows necessity, $k=500$ shows Top-$k$ restoration above random, and $k=d/2$ is largely capacity-driven. Partial-occlusion sweeps reveal that high-overlap coordinate-token sets can hurt at $k=1000$ and help mainly at half-width, while population corruption-derived sets provide no reliable fixed repair set. Edge-attribution patching shows coordinate-token paths are high precision but low recall for detection recovery, and RMSNorm quasi-layer controls do not close the endpoint-repair gap. Endpoint coordinate triage is therefore a useful circuit prior, but occlusion recovery requires a separate benchmark.

补充信息

↑