发表机构
The University of Melbourne; Information Technology University; Westlake University(墨尔本大学; 信息技术大学; 西湖大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出InstructMixup数据增强方法,利用轻量级检测器提取显著补丁,经指令引导生成模型细化后融合回原样本,还注入分形结构增加多样性,推导二阶近似,实验表明该方法在多基准测试中优于多种竞争方法。
AI 中文摘要
在图像和视频技术中,数据增强广泛用于提高深度视觉模型的泛化能力,基于样本插值的混合策略成为主流方法。然而,计算信息丰富的混合区域会增加大量开销,跨不同图像混合内容常破坏结果样本的语义完整性。我们提出InstructMixup,一种在单个视觉样本内构建具有挑战性且标签一致的训练样本的数据增强方法。它先使用轻量级显著性检测器从样本中提取多尺度显著补丁,用指令引导生成模型细化每个补丁,再将编辑后的补丁融合回同一样本的非显著区域,此步骤训练成本可忽略不计。为使学习表示更多样化,它以自适应比例在相同显著区域注入自相似分形结构。我们推导了所得邻域风险的二阶近似,表明该方法同时强制对生成编辑的不变性并抑制沿扰动显著方向的曲率,且通过实验验证了这两个预测。我们在涵盖粗粒度和细粒度分类、对损坏和遮挡的鲁棒性、校准以及迁移和自监督学习的七个基准上,对从小型到大型的骨干网络(如卷积神经网络、视觉Transformer和视觉语言基础模型)进行评估,InstructMixup优于九种竞争增强方法,在所有基准上超过最强基线。
英文摘要
In image and video technologies, data augmentation is widely used to improve the generalization of deep visual models, and mixup-based strategies that interpolate between samples have become the dominant approach. However, computing informative mixing regions adds substantial overhead, and blending content across different images frequently disrupts the semantic integrity of the resulting sample. We propose \our{}, a data augmentation method that constructs challenging yet label-consistent training samples entirely within a single visual sample. \our{} first extracts multi-scale salient patches from the sample using a lightweight saliency detector, refines each patch with an instruction-guided generative model, and blends the edited patch back into the non-salient regions of the same sample; because the generative edits are computed once and cached offline, this step adds negligible training cost. To further diversify the learned representation, \our{} injects self-similar fractal structure into the same salient regions at an adaptive ratio, so each training sample carries both fractal and non-fractal structure. We derive a second-order approximation of the resulting vicinal risk, showing that the method simultaneously enforces invariance to the generative edit and suppresses curvature along the perturbed salient directions, and we verify both predictions empirically. We evaluate on small to large backbones for instance Convolutional Neural Networks (CNNs), Vision Transformers (ViTs) and Vision-Language Foundational Models (VLMs) across seven benchmarks covering coarse- and fine-grained classification, robustness to corruption and occlusion, calibration, and transfer and self-supervised learning, InstructMixup outperforms nine competing augmentation methods, surpassing the strongest baseline across all benchmarks.