发表机构
University of Gujrat; Khalifa University(古吉拉特大学; 哈利法大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过跨水下、空中、放射学三个物理域的实验,重新评估BLIP微调中梯度范数不平衡的纠正方法,发现降低不平衡并不一致地提升性能,且纠正LoRA配置可提高BLEU-4。
AI 中文摘要
视觉-语言模型的视觉通路和语言通路之间的梯度幅度不平衡通常被视为需要纠正的缺陷。我们针对一类纠正方法检验了这一前提,并特意排除了自适应、信号驱动的方案(如BalGrad、OGM、PMR、CGGM),这些方案属于机制上不同的类别,不在本研究范围内。我们测量了语言到视觉的梯度范数比(以参数归一化形式报告),在九个微调条件、三个随机种子和三个跨越不同物理域偏移(水下、空中、放射学)的字幕数据集上,发现不平衡幅度在不同域之间差异显著,且没有可预测的顺序。单纯降低学习率可大幅减少不平衡,并在每个数据集上达到与最佳方法相差几个BLEU点的水平。分阶段冻结在每个域上都降低了该比率,但从未排名第一;仅调整调度表的对照实验在一个数据集上隔离出冻结是原因,但在另外两个数据集上则不然。强制两个梯度组幅度相等,使每个域上的每参数不平衡接近零,但在不同数据集上、相同设置下,这既是研究中最好的结果,也是全微调方法中最差的排名。梯度范数比的降低并不能跨域一致地预测字幕性能,达到给定平衡水平的方式与水平本身同样重要。作为次要发现,一个常用于BLIP的LoRA配置会静默地适应零个视觉参数;纠正该问题可在所有三个数据集上提高BLEU-4。
英文摘要
Imbalanced gradient magnitudes between the visual and language pathways of a vision-language model are often treated as a defect to be corrected. We test that premise for one family of correction, deliberately excluding adaptive, signal-driven schemes (e.g. BalGrad, OGM, PMR, CGGM), which are a mechanistically distinct class outside this study's scope. Measuring the language-to-visual gradient-norm ratio, reported in parameter-normalised form, across nine fine-tuning conditions, three seeds, and three captioning datasets spanning distinct physical domain shifts -- underwater, aerial, radiological -- we find imbalance magnitude varies markedly across domains with no predictable ordering. A plain learning-rate reduction cuts imbalance substantially and lands within a few BLEU points of the best method on every dataset. Staged freezing reduces the ratio on every domain yet never ranks first; a schedule-only control isolates freezing as the cause on one dataset but not the other two. Forcing the two gradient groups to equal magnitude drives per-parameter imbalance close to zero on every domain, yet is both the best result in the study and the worst placement among full fine-tuning methods, on different datasets, with identical settings. Reductions in gradient-norm ratio do not consistently predict captioning performance across domains, and how a given level of balance is reached matters as much as the level itself. As a secondary finding, a commonly reused LoRA configuration applied to BLIP silently adapts zero visual parameters; correcting it improves BLEU-4 on all three datasets.
Comments10 pages, 14 figures. Supplementary material included as ancillary file