发表机构
School of Artificial Intelligence, University of Chinese Academy of Sciences; MAIS, Institute of Automation, Chinese Academy of Sciences; Alibaba Token Hub, Alibaba Group(中国科学院大学人工智能学院; 中国科学院自动化研究所; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于数据内在一致性(DIC)的自适应视觉指令选择方法DICS,仅用25%数据就超越现有最优方法,还构建600万样本语料库,用不足25%训练数据达到InternVL3-8B-Instruct的94.52%性能
AI 中文摘要
视觉指令微调对提升视觉语言模型(VLMs)的视觉语言对齐和指令跟随能力至关重要,但从快速扩充的数据集中在固定比例约束下筛选最优子集仍是重大瓶颈。现有方法多依赖分布多样性或启发式过滤,常忽略单个样本内部的连贯性。为填补这一空白,我们提出数据内在一致性(DIC),这是一种用于量化样本级组件间一致性的自评分指标。DIC包含两个模块:视觉信息一致性(VIC),评估视觉内容与指令的对齐度;响应信息一致性(RIC),评估响应相对于指令的连贯性。基于DIC,我们引入数据内在一致性选择(DICS),这是一种自适应数据选择方法,可在不同数据预算下优化高样本内一致性与全局分布多样性的权衡。大量实验表明,DICS在不同数据集规模和模型架构上均持续优于现有最优方法,仅使用LLaVA-1.5-665K数据的25%,就超过了全数据集微调的效果。我们进一步整理出DICS-6M,这是一个包含600万样本的多模态指令语料库,支撑了迄今为止最大规模的视觉指令选择研究;值得注意的是,DICS使用不到官方InternVL3-8B-Instruct报告训练数据的25%,就达到了其94.52%的性能。代码可在该https URL查看
英文摘要
Visual instruction tuning is crucial for advancing the vision-language alignment and instruction-following capabilities of Vision-Language Models (VLMs). However, identifying optimal subsets under a fixed ratio constraint from rapidly expanding datasets remains a significant bottleneck. While existing methods largely depend on distribution diversity or heuristic filtering, they often overlook the internal coherence within individual samples. To bridge this gap, we propose Data Intrinsic Consistency (DIC), a self-scoring metric designed to quantify the sample-level inter-component consistency. DIC consists of two modules: Visual Information Consistency (VIC), evaluating the alignment between visual content and instructions, and Response Information Consistency (RIC), assessing response coherence relative to the instruction. Building upon DIC, we introduce Data Intrinsic Consistency Selection (DICS), an adaptive data selection method that optimizes the trade-off between high intra-sample consistency and global distributional diversity under varying data budgets. Extensive experiments demonstrate that DICS consistently outperforms state-of-the-art methods across diverse dataset scales and model architectures, surpassing full-dataset fine-tuning while using only 25% of the LLaVA-1.5-665K data. We further curate DICS-6M, a 6M-sample multi-modal instruction corpus that enables the largest-scale visual instruction selection study to date; remarkably, DICS reaches 94.52\% of the official InternVL3-8B-Instruct performance using less than 25\% of its reported training data. Code can be seen at https://github.com/cqu-student/DICS
CommentsAccepted by EMNLP2026 Findings