arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32529cs.CV

SetOPD:从少量视觉示例到多模态候选集用于遥感开放提示检测

SetOPD: From Few Visual Exemplars to Multimodal Candidate Sets for Remote-Sensing Open-Prompt Detection

Jinlong Hu, Yi Zhang, Zhiqi Xia, Yikang Zhou, Shunping Ji

AI总结:

SetOPD提出从集合视角改进遥感开放提示检测:通过BR提示保留视觉示例细节,并在候选状态层面用PQA进行显式多模态协作,解决现有方法对视觉模态利用不足的问题。

AI中文摘要:

开放提示检测器允许用户通过文本、视觉示例或两者来指定目标。我们认为现有设计在两个方面未充分利用视觉模态。首先,多个示例通常被压缩为单个类别级嵌入。这使视觉提示文本化:所得向量扮演另一个类别名称的角色,通常与文本路径对齐或注入其中,并且在仅有少量异质示例时可能不是最优的。其次,现有方法主要在提示或表示空间中进行交互,此时模态特定的检测状态尚未形成。我们从集合的角度解决这两个问题:SetOPD从共享的提示条件初始化中保留模态特定的解码状态,并在候选状态层面执行显式的多模态协作。针对第一个问题,我们引入BR提示,它读取每个框选示例的完整场景上下文,并将汇集的支持证据分解为基锚和学习的残差校正;所得提示的容量固定,与示例数量无关,并驱动其自身的视觉检测路径。针对第二个问题,我们将文本-视觉协作从表示级融合重新定义为候选集建模问题。配对查询仲裁(PQA)仅在模态特定候选状态形成后执行显式的跨模态状态仲裁。两个读取器共享查询初始化,因此它们的候选按索引配对;一个学习的门控在每个配对内进行仲裁,随后一个置换等变模块对融合后的集合进行推理。

英文摘要:

Open-prompt detectors allow users to specify targets with text, visual exemplars, or both. We argue that existing designs underuse the visual modality in two ways. First, multiple exemplars are commonly compressed into a single class-level embedding. This textualizes visual prompting: the resulting vector plays the role of another class name, is often aligned to or injected into the text pathway, and may be suboptimal when only a few heterogeneous exemplars are available. Second, existing methods interact primarily in prompt or representation space, before modality-specific detection states are formed. We address both issues from a set perspective: \setopd preserves modality-specific decoding states from a shared prompt-conditioned initialization and performs explicit multimodal collaboration at the candidate-state level. For the first issue, we introduce \br prompting, which reads every boxed exemplar in its full scene context and decomposes the pooled support evidence into a Base anchor and a learned Residual correction; the resulting prompt has fixed capacity regardless of the number of examples and drives its own visual detection pathway. For the second, we recast text--visual collaboration from representation-level fusion into a candidate-set modeling problem. Paired-Query Arbitration (\pqa) then performs explicit cross-modal state arbitration only after modality-specific candidate states have been formed. The two readers share query initialization so their candidates are paired by index; a learned gate arbitrates within each pair, followed by a permutation-equivariant module that reasons over the fused set.

↑