arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

门总是关闭:关于向冻结的视觉语言模型注入辅助信号

The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models

Moshiur Farazi, Sameera Ramasinghe, Bekir Sait Ciftler, Mahbub Ahmed Turza, Shafin Rahman

arXiv 2607.23335首次发表:更新:

发表机构

University of Doha for Science and Technology; Pluralis Research; North South University(多哈科技大学; 普卢拉利斯研究公司; 南北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言模型中辅助信号注入问题,发现优化器常使门控路径关闭。利用抑制现象,通过双曲视觉关系图的几何辅助损失正则化LoRA微调,实验表明该方法能保持关系准确率,还指出嵌入范数对齐及固定注入的必要性。

AI 中文摘要

在视觉语言模型(VLMs)中,辅助信号路径通常配备可学习的门,以便优化器决定允许多少信号通过。研究发现优化器几乎总是决定为零:在五种注入设计中,每个有门控的路径在行为上都关闭了,即使门参数名义上会通过30 - 45%的信号,在推理时消融该路径,准确率也不变。这种抑制现象归因于两种情况,一种是通过图像衍生信号的字幕不变性形式化的死梯度情况,另一种是辅助信号会积极损害损失的负效用情况。利用这种抑制现象,通过来自双曲视觉关系图的几何辅助损失(IoA驱动的蕴含锥和洛伦兹流形上的角度排斥)对LoRA微调进行正则化,仅在训练时通过前向传播耦合,推理时丢弃。按问题类型分解GQA显示出清晰的分离。三种在推理时没有几何损失的配置在关系问题上损失2.85 - 3.39个百分点,而在属性问题上增加约1个百分点;第四种配置在训练时使用这些损失但通过软提示进行推理,在关系问题上损失5.14个百分点,在属性问题上仅增加0.23个百分点,所以仅训练时的正则化在没有几何推理路径的情况下无法保护关系准确率。在推理时保留几何路径的配置保持了普通级别的关系准确率并与属性增益相匹配。在VSR的分布外,RMS前缀方法保留了空间信号;去除几何损失(G2)使VSR下降4.6个百分点,将它们分离为分布外的来源。第二个结果是:嵌入范数对齐对于生成安全的前缀注入是必要的,并且可学习的门应该被固定的、非可选的匹配尺度注入所取代。

英文摘要

Auxiliary signal pathways in VLMs are routinely fitted with learnable gates so the optimiser can decide how much of the signal to admit. We find that the optimiser almost always decides on zero: across five injection designs, every gated pathway becomes behaviourally closed, with accuracy invariant to ablating the pathway at inference even when the gate parameter would nominally pass 30-45% of the signal. We attribute this suppression phenomenon to two regimes, a dead-gradient regime formalised through the caption-invariance of image-derived signals, and a negative-utility regime in which the auxiliary signal actively hurts the loss. Rather than fight suppression, we exploit it: we regularise LoRA fine-tuning with geometric auxiliary losses from hyperbolic visual relational graphs (IoA-driven entailment cones and angular repulsion on the Lorentz manifold), coupled only through the forward pass at training time and dropped at inference. Disaggregating GQA by question type exposes a clean dissociation. Three configurations without geometric losses at inference lose 2.85-3.39pp on relational questions while gaining ~1pp on attribute questions; a fourth that trains with the losses but infers through a soft prompt loses 5.14pp on rel for only +0.23pp on attr, so training-time regularisation alone does not protect relational accuracy without a geometric inference pathway. Configurations that keep the geometric pathway at inference preserve vanilla-level relational accuracy and match the attribute gain. Out of distribution on VSR, the RMS-prefix recipe preserves the spatial signal; stripping the geometric losses (G2) collapses VSR by 4.6pp, isolating them as the OOD source. A secondary result: embedding-norm alignment is necessary for generation-safe prefix injection, and learnable gates should be replaced with fixed, non-optional injection at matched scales.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑