arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22584cs.AI

用于鲁棒的两阶段以对象为中心的视觉推理的弱监督概念瓶颈学习

Weakly supervised concept Bottleneck Learning for Robust Two stage Object centric visual reasoning

发表机构吕贝克大学 · 乌尔姆大学 · 班贝格大学
查看机构详情
  • University of Lübeck(吕贝克大学)
  • Ulm University(乌尔姆大学)
  • University of Bamberg(班贝格大学)

机构由 AI 辅助整理,请以论文原文为准。

Sparsh Tiwari, Gesina Schwalbe, Bettina Finzel

首次发表
浏览论文内容

中文总结 AI 辅助

针对两阶段视觉推理需昂贵标注的问题,提出动态正交概念瓶颈(D-OCB)框架,在极弱监督下提取对齐人类的符号谓词,通过动态超参数分配、子空间相关性惩罚及动态维度分配,实现高概念对齐与下游推理准确率,性能媲美或优于端到端范式。

中文摘要 AI 辅助

两阶段神经符号架构为视觉问题解决提供了一种优雅的范式,它将预定义符号的联结主义感知与之后可能进行的关系推理清晰地分离开来。然而,将高层谓词锚定到视觉框架通常需要获取成本高昂的标注。在这项工作中,我们引入了动态正交概念瓶颈(Dynamic Orthogonal Concept Bottleneck,D-OCB),这是一种以对象为中心的槽变分自编码器(slot-VAE)框架,旨在在极弱监督下提取与人类对齐的符号谓词。D-OCB通过在训练过程中动态学习最优超参数分配,消除了对损失平衡系数的繁琐手动调整。为了注入概念类别独立性的先验知识,除了标准的重建自监督外,我们还对概念子空间之间的相关性进行惩罚。至关重要的是,为了应对极低监督 regime 的不稳定性,D-OCB 融入了动态维度分配机制;这种自适应公式化方法使表现良好的概念能够将潜在维度让渡给表现滞后的概念,有效防止表示崩溃并显著提高整体概念准确率。通过广泛的实证评估,我们证明我们的框架使用最少的标签预算即可实现高概念对齐度和下游视觉推理准确率,与端到端范式相当或优于后者。

英文摘要

Two-stage neuro-symbolic architectures provide an elegant paradigm for visual problem solving by cleanly separating connectionist perception of predefined symbols from possibly later defined relational reasoning thereon. However, anchoring high-level predicates into visual frames typically necessitates annotations that are expensive to acquire. In this work, we introduce the Dynamic Orthogonal Concept Bottleneck (D-OCB), an object-centric slot- VAE framework designed to extract human-aligned symbolic predicates under extremely weak supervision. D-OCB eliminates the arduous manual tuning of loss-balancing coef- ficients by dynamically learning optimal hyperparameter allocations during training. To infuse prior knowledge on independence of concept categories, in addition to standard re- construction self-supervision we penalize correlation across concept subspaces. Crucially, to combat the instability of very low supervision regimes, D-OCB incorporates a dynamic di- mensionality allocation mechanism; this adaptive formulation allows well-represented con- cepts to yield latent dimensions to underperforming concepts that are lagging behind, effectively preventing representation collapse and significantly improving overall concept accuracy. Through an extensive empirical evaluation, we demonstrate that our framework achieves high concept alignment and downstream visual reasoning accuracy using minimal label budgets, matching or outperforming end-to-end paradigms.

↑