arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10292cs.CV

各向同性嵌入扰动用于鲁棒视觉语言编码器

Isotropic Embedding Perturbations for Robust Vision Language Encoders

  • Soongsil Univ.(崇实大学)
  • NAVER AI Lab(NAVER AI实验室)
  • Ewha W. Univ.(梨花女子大学)

机构由 AI 辅助整理,请以论文原文为准。

Hyesong Choi, Daeun Kim, Song Park, Taekyung Kim, Byeongho Heo, Sangdoo Yun, Dongbo Min, Dongyoon Han

AI总结:

Aether是一种即插即用的嵌入空间各向同性扰动增强方法,通过受控alpha混合施加扩散式随机扰动,在多种架构和识别任务上超越传统像素级增强组合,显著提升视觉语言编码器的多模态对齐性能。

AI中文摘要:

数据增强是训练现代深度视觉和多模态模型的基础。虽然诸如RandAug、CutMix、Mixup、RandErase和DropPath等单独方法提供了强大的正则化效果,但由于功能重叠,它们的组合使用在性能上已趋于饱和,且激进的像素级操作可能破坏精细的跨模态对齐。这种饱和促使人们在嵌入空间而非输入空间中寻找新的增强轴。我们引入了Aether,一种简单的即插即用方法,通过受控的alpha混合在嵌入空间中应用扩散式随机扰动,专门设计用于提供保持语义一致性的各向同性正则化。受语言模型中特征空间扰动和生成预训练中图像退化现象的启发,Aether引入温和而有效的扰动,平滑表示而不损害强视觉语言编码器所需的细粒度结构信息。在多种架构和多个识别任务中,Aether相较于结合CutMix、Mixup、DropPath和RandAug的高级配方持续带来性能提升——这种改进程度在现代增强替代方案中很少见。值得注意的是,Aether在多模态对齐方面展现出卓越的有效性,在传统像素空间增强失效的场景中取得成功,通过提供稳定、各向同性的正则化信号,尊重高维特征空间的完整性。

英文摘要:

Data augmentation is fundamental to training modern deep vision and multimodal models. While individual methods, such as RandAug, CutMix, Mixup, RandErase, and DropPath, offer strong regularization effects, their combined use has saturated in performance due to overlapping functionalities, and aggressive pixel-level manipulations may disrupt delicate cross-modal alignment. This saturation motivates the search for a new augmentation axis within the embedding space rather than the input space. We introduce Aether, a simple plug-in method that applies diffusion-style random perturbations in the embedding space via controlled alpha-mixing, specifically designed to provide isotropic regularization that remains semantically consistent. Inspired by feature-space perturbations in language models and image degradation in generative pretraining, Aether induces mild yet effective perturbations that smooth the representations without compromising the fine-grained structural information required for strong vision-language encoders. Across diverse architectures and across multiple recognition tasks, Aether delivers consistent gains over the advanced recipe combining CutMix, Mixup, DropPath, and RandAug---a level of improvement rarely observed with modern augmentation alternatives. Notably, Aether demonstrates superior effectiveness in multi-modal alignment, succeeding where traditional pixel-space augmentations fail by providing a stable, isotropic regularization signal that respects the integrity of the high-dimensional feature space.

补充信息

↑