DecoupleMix:用于可扩展视觉语言模型数据配方的解耦比率搜索和凸分配
DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes
浏览论文内容
中文总结 AI 辅助
该研究针对视觉语言模型数据构建缺乏原则性标准的问题,提出DecoupleMix框架,将其解耦为类间比率搜索和类内凸分配两个子问题,实验证明该方法优于启发式基线,且最优比率可跨规模转移,提升了VLM竞争力。
中文摘要 AI 辅助
虽然视觉语言模型(VLM)的数据整理越来越活跃,但构建预训练混合数据的公共实践在很大程度上仍然是启发式的:从业者堆叠通过质量过滤的数据集,凭直觉设置跨域比率,并且缺乏纳入新数据的原则性、可归因标准,同时前沿配方仍未公开。我们将数据构建制定为一个系统的混合优化问题,并通过将混合解耦为两个正交子问题,将其转变为一门可重复的工程学科:跨能力的类间比率和类别内的类内比率。对于类间分配,我们使用单变量迭代搜索;对于类内组成,我们应用多维、数据集级评估来评分质量和难度,并将选择制定为具有多样性目标的约束凸优化。DecoupleMix框架提供了两个关键能力:指导接下来收集什么数据,并使数据集验证成为一个可控、可归因的实验。实验表明我们的方法始终优于启发式基线。此外,在小规模代理上发现的最优比率无需重新调整即可无缝转移到更大规模。使用80B额外的多模态继续预训练令牌,我们的VLM与使用大量多模态预算训练的强大开源模型具有竞争力。
英文摘要
While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios by intuition, and lack a principled, attributable criterion for admitting new data, while frontier recipes remain undisclosed. We formulate data construction as a systematic mixture-optimization problem and turn it into a reproducible engineering discipline by decoupling the mixture into two orthogonal sub-problems: inter-class ratios across capabilities and intra-class ratios within a category. For inter-class allocation, we use a single-variable iterative search; for intra-class composition, we apply a multidimensional, dataset-level assessment scoring Quality and Difficulty, and formulate selection as a constrained convex optimization with a diversity objective. The DecoupleMix framework delivers two critical capabilities: guiding what data to collect next and rendering dataset validation a controlled, attributable experiment. Experiments show our approach consistently surpasses heuristic baselines. Moreover, optimal ratios discovered on small-scale proxies transfer seamlessly to larger scales without retuning. Using 80B additional multimodal continue-pretraining tokens, our VLM is competitive with strong open-source models trained with substantially larger multimodal budgets.
发表机构
- ByteDance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。