arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

条件视觉证据效用:冻结的视觉-语言编码器中的状态依赖排名反转

Conditional Visual Evidence Utility: State-Dependent Rank Reversals in Frozen Vision-Language Encoders

Yunxuan Fang, Xinhe Wang

arXiv 2608.28316首次发表:更新:

发表机构

Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现冻结的OpenCLIP和SigLIP在受控组合视觉搜索中存在状态依赖的视觉证据排名反转,更新证据排序可保留跨评估者的决策价值,推动条件性评估视觉-语言证据使用。

AI 中文摘要

静态重要性分数将视觉证据压缩为单一排名,但在观察到一个线索后,剩余证据的价值可能发生变化。我们在受控的组合视觉搜索中研究这种可能性,其中颜色、形状和纹理证据可被独立呈现,并在不同获取状态下测量它们的条件边际效用。在对800个场景的保留验证中,冻结的OpenCLIP和SigLIP表现出稳健的状态依赖排名反转,这些反转集中在旨在诱导排序变化的候选重叠区域。该结构在两种证据积累构建方式和十种等效查询表述下均保持存在,但在查询-场景打乱时消失。我们还探究这些反转是否对决策有影响。在验证后的探索性匹配首动作分析中,仅在第一次获取后重新排序,当决策在一种证据模式、表述或主干下选择并在另一种下评估时,会产生正的第二步效用。这些结果共同表明,在该受控设置中,证据重要性是状态依赖的,更新证据排序可在评估者变化时保留与决策相关的价值。它们推动人们对视觉-语言证据的使用进行条件评估,而非通过单一静态排名,同时为未来的自适应证据选择方法提供了可测量的目标。

英文摘要

Static importance scores compress visual evidence into a single ranking, but the value of remaining evidence can change after one cue has been observed. We study this possibility in controlled compositional visual search, where color, shape, and texture evidence can be independently exposed and their conditional marginal utility measured across acquisition states. In a held-out confirmation on 800 scenes, frozen OpenCLIP and SigLIP exhibit robust state-dependent rank reversals that concentrate in candidate-overlap regimes designed to induce ordering changes, persist across two evidence-accumulation constructions and ten equivalent query wordings, and collapse to near-chance-scale behavior under query-scene derangement. A subsequent role-balanced follow-up on 1,200 scenes rotates the abstract roles of initially strong, redundancy-inducing, and comparator attributes; the positive-minus-negative reversal contrast remains positive across all 24 role-permutation, backbone, and evidence-mode cells, although residual attribute-identity effects remain. We further distinguish measured replanning opportunity from prospective predictability. Matched-first-action utility analyses show substantial opportunity to rerank remaining evidence, but lightweight predictors using posterior-based or acquired-embedding state representations do not establish a robust incremental advantage of acquired-state information over legal static controls on the role-balanced benchmark. Together, these results show that conditional visual evidence utility is reliably state dependent in this controlled setting, while separating the existence of changing utility from the stronger claim that those changes are prospectively predictable by a learned selector.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑