DVLA-RL++:用于小样本学习的带强化学习门控的双层级视觉-语言对齐方法
DVLA-RL++: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
浏览论文内容
中文总结 AI 辅助
本文针对小样本学习中支持原型易受上下文污染的问题,提出带互补语义纯化与反事实强化学习门控的DVLA-RL++,在多类基准上较DVLA-RL平均提升1.4%准确率,达到当前最优性能。
中文摘要 AI 辅助
小样本学习旨在从有限的标注示例中识别新类别。近期研究引入文本语义以弥补有限的视觉观测,优化类别表示。然而,图像-文本的高一致性可能同时反映物体固有属性与偶然上下文,导致支持原型易受上下文污染。为解决该问题,本文提出DVLA-RL++,它在DVLA-RL基础上新增互补语义纯化(CSP)与反事实强化学习门控(CRG)。具体而言,CSP从标注支持样本生成固有描述与干扰描述,并比较二者与各支持 token 的一致性;依赖歧义度的拒绝边际引导稀疏证据分配,而固有语义锚点会填充未分配的质量,在视觉证据不可靠时提供弃权(不执行)的备选方案。CRG利用平衡识别性能与干扰暴露的奖励,学习层级语义融合强度,同一 episode 中独立执行的参考轨迹提供配对学习信号。理论分析将保留的证据与锚点质量关联至原型稳定性,并建立无偏在线策略梯度估计的条件。在标准、细粒度及跨域基准上的实验显示,该方法达到了当前最优准确率,较DVLA-RL平均提升1.4%。项目页面可通过此https URL访问。
英文摘要
Few-shot learning aims to recognize novel categories from limited labeled examples. Recent studies incorporate textual semantics to compensate for limited visual observations and improve class representations. However, high image-text agreement may reflect both intrinsic object properties and incidental context, making support prototypes susceptible to contextual contamination. To address this problem, we propose DVLA-RL++, which extends DVLA-RL with complementary semantic purification (CSP) and counterfactual reinforcement-learning gating (CRG). Specifically, CSP generates intrinsic and nuisance descriptions from labeled supports and compares their agreement with each support token. An ambiguity-dependent rejection margin guides sparse evidence allocation, while an intrinsic semantic anchor fills the unassigned mass to provide a fallback when visual evidence is unreliable. CRG learns layer-wise semantic fusion strengths using a reward that balances recognition performance and nuisance exposure. An independently executed reference trajectory on the same episode provides a paired learning signal. Theoretical analysis relates retained evidence and anchor quality to prototype stability and establishes conditions for unbiased on-policy gradient estimation. Experiments on standard, fine-grained, and cross-domain benchmarks show state-of-the-art accuracy, with an average gain of 1.4% over DVLA-RL. The project page is available at https://peacelwh.github.io/TPAMI27-DVLA-RLpp/.
发表机构
- School of Software, Shandong University(山东大学软件学院)
- Shenzhen Loop Area Institute(深圳河套学院)
- School of Computer and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)
- School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。