SEER:用于受控空间关系分类的自 grounding 证据接口
SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification
AI总结:
本文提出针对冻结VLM的无训练推理时证据接口SEER,通过构建查询特定证据缓解空间关系分类错误,在多数据集上取得显著性能提升,证明该干预措施的有效性。
AI中文摘要:
空间关系问题要求模型在比较布局前识别查询的主体和客体,然而视觉语言模型(VLM)即使能识别两个实体,仍可能因选择错误实例或模糊全局视图而答错。本文探究将查询特定证据显性化能否缓解该问题,提出SEER(用于实体-关系推理的自 grounding 证据),这是一种针对冻结VLM的无训练推理时证据接口。SEER在配对定位时隐藏候选关系,构建带有显性主体/客体角色的查询特定视图,同时保留完整图像和稀疏框几何作为补充证据。对于具有精确反向支持的关系选择协议,可选的精调会交换实体角色,仅当恰好一个视觉状态服从对应反向关系时才改变前向决策。在图像不相交的GQA-Train900测试集(模型评分前冻结)上,SEER相比Full提升+3.94 [2.17,5.72];在标签独立的 grounding 顺序平衡下,以及实体名称唯一的535行数据上,该提升仍为正。在三个模型的全部2434个过滤后的EmbSpatial配对-关系问题上,未改变的协议提升幅度为+4.35至+11.79。匹配对照实验将局部重聚焦与角色显性条件分离,结果表明查询特定证据构建是主要干预措施,而互逆一致性是较小的协议特定精调。
英文摘要:
Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both entities and still answer from the wrong instance or an ambiguous global view. We ask whether making query-specific evidence explicit can mitigate this failure and propose SEER (Self-grounded Evidence for Entity-Relation Reasoning), a training-free inference-time evidence interface for frozen VLMs. SEER hides candidate relations during pair localization, constructs a query-specific view with explicit subject/object roles, and retains the full image and sparse box geometry as complementary evidence. For relation-choice protocols with exact inverse support, an optional refinement swaps the entity roles and changes the forward decision only when exactly one visual state obeys the corresponding inverse relation. On an image-disjoint GQA-Train900 test frozen before model scoring, SEER pools to +3.94 [2.17,5.72] over Full; the gain remains positive under label-independent grounding-order counterbalancing and on the 535 rows whose entity names are unique. The unchanged protocol yields +4.35 to +11.79 on all 2,434 filtered EmbSpatial pair-relation questions across three models. Matched controls separate local refocus from role-explicit conditioning. These results establish query-specific evidence construction as the principal intervention, with reciprocal consistency as a smaller protocol-specific refinement.