arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12095cs.CVcs.LG

DVLA-RL++:用于小样本学习的带强化学习门控的双层级视觉-语言对齐方法

DVLA-RL++: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning

Wenhao Li, Xianjing Meng, Qiangchang Wang, Zhongyi Han, Yilong Yin, Liqiang Nie

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对小样本学习中支持原型易受上下文污染的问题,提出带互补语义纯化与反事实强化学习门控的DVLA-RL++,在多类基准上较DVLA-RL平均提升1.4%准确率,达到当前最优性能。

中文摘要 AI 辅助

小样本学习旨在从有限的标注示例中识别新类别。近期研究引入文本语义以弥补有限的视觉观测,优化类别表示。然而,图像-文本的高一致性可能同时反映物体固有属性与偶然上下文,导致支持原型易受上下文污染。为解决该问题,本文提出DVLA-RL++,它在DVLA-RL基础上新增互补语义纯化(CSP)与反事实强化学习门控(CRG)。具体而言,CSP从标注支持样本生成固有描述与干扰描述,并比较二者与各支持 token 的一致性;依赖歧义度的拒绝边际引导稀疏证据分配,而固有语义锚点会填充未分配的质量,在视觉证据不可靠时提供弃权(不执行)的备选方案。CRG利用平衡识别性能与干扰暴露的奖励,学习层级语义融合强度,同一 episode 中独立执行的参考轨迹提供配对学习信号。理论分析将保留的证据与锚点质量关联至原型稳定性,并建立无偏在线策略梯度估计的条件。在标准、细粒度及跨域基准上的实验显示,该方法达到了当前最优准确率,较DVLA-RL平均提升1.4%。项目页面可通过此https URL访问。

英文摘要

Few-shot learning aims to recognize novel categories from limited labeled examples. Recent studies incorporate textual semantics to compensate for limited visual observations and improve class representations. However, high image-text agreement may reflect both intrinsic object properties and incidental context, making support prototypes susceptible to contextual contamination. To address this problem, we propose DVLA-RL++, which extends DVLA-RL with complementary semantic purification (CSP) and counterfactual reinforcement-learning gating (CRG). Specifically, CSP generates intrinsic and nuisance descriptions from labeled supports and compares their agreement with each support token. An ambiguity-dependent rejection margin guides sparse evidence allocation, while an intrinsic semantic anchor fills the unassigned mass to provide a fallback when visual evidence is unreliable. CRG learns layer-wise semantic fusion strengths using a reward that balances recognition performance and nuisance exposure. An independently executed reference trajectory on the same episode provides a paired learning signal. Theoretical analysis relates retained evidence and anchor quality to prototype stability and establishes conditions for unbiased on-policy gradient estimation. Experiments on standard, fine-grained, and cross-domain benchmarks show state-of-the-art accuracy, with an average gain of 1.4% over DVLA-RL. The project page is available at https://peacelwh.github.io/TPAMI27-DVLA-RLpp/.

发表机构

  • School of Software, Shandong University(山东大学软件学院)
  • Shenzhen Loop Area Institute(深圳河套学院)
  • School of Computer and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)
  • School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑