发表机构
Changwon National University; Chung-Ang University(昌原大学; 中央大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出无答案密钥的协议,探究冻结VLM对无图像目标标记编辑的响应规律,在两个遥感数据集和三个冻结LM主干上验证了相关结构,发布了相关资源。
AI 中文摘要
用VLM回答关于场景的假设性问题,通常是将假设作为文本注入,或用生成模型重新绘制场景。相反,我们将编辑操作移至模型输入前的表示层面,图像被抽象为一组目标级标记,且原始图像不会进入VLM。该设计基于一个开放问题:冻结VLM何时会对这类标记编辑做出实际响应?我们提出一种无答案密钥的协议:不对编辑后的答案进行标注,它对答案在逻辑上确定的编辑进行评分,并通过反转每个可评分的选择来自我审计。该协议揭示了三种结构:响应并非自由产生,在所有三种操作下,是显式的编辑教学而非普通的VQA训练,在密集场景中产生响应,在稀疏场景中使响应倍数增加;一旦响应产生,其由标记的清晰度和密度决定,可部署的检测器+分割器标记与 oracle 具有竞争力,且在VRSBench上优于oracle;阅读是一个可分离的维度:无图像的标记路径保留了匹配的 patch 标记基线的92%-96%自由文本VQA性能,且答案明显依赖于标记。响应、清晰度和阅读结构在两个遥感数据集(iSAID、VRSBench)和三个冻结LM主干中保持符号一致性。我们发布了探测生成器、记录、判断日志和代码。
英文摘要
Answering what-if queries about a scene with a VLM usually means injecting the assumption as text or repainting the scene with a generative model. We instead move the edit to the representation level, before the model input. The image is abstracted into a set of object-level tokens, and the original image never enters the VLM. This design rests on an open question: when do frozen VLMs actually respond to such token edits? We introduce an answer-key-free protocol: no post-edit answer is annotated. It scores edits whose answers are logically determined, and audits itself by reversing each scoreable choice. The protocol reveals three structures. The response is not free: explicit edit teaching, not ordinary VQA training, produces it in dense scenes and multiplies it in sparse ones, on all three operations. Once on, it is governed by token cleanliness and density, with deployable detector+segmenter tokens competitive with the oracle and outperforming it on VRSBench. And reading is a separable axis: the image-free token route preserves 92-96% of a matched patch-token baseline's free-text VQA, and the answers measurably depend on the tokens. The response, cleanliness, and reading structures are sign-preserved across two remote-sensing datasets (iSAID, VRSBench) and three frozen LM backbones. We release the probe generator, records, judge logs, and code.
Comments10 pages, 3 figures, 5 tables