配对多模态缩放定律
Paired Multimodal Scaling Laws
浏览论文内容
中文总结 AI 辅助
本研究提出配对多模态缩放定律,将损失分解为四个幂律通道,揭示配对数据对协同损失的临界阈值影响,并在实验中以3.2%误差优于现有定律。
中文摘要 AI 辅助
现有的多模态缩放定律在测试后凭经验拟合多模态项,并且从未在固定数据预算下改变多模态配对数据的数量。我们研究了在每种模态总数据量相同的情况下,改变配对数据数量如何影响多模态分类任务中的损失曲线。我们在三种不同环境中训练模型,并运行改变数据规模和配对预算的实验扫描。配对比例对损失有显著影响,且这种影响直接与任务包含的信息协同量相关。只有配对数据能够减少协同损失,而未配对数据可以减少冗余或单模态信息,直至单模态下限。与传统缩放定律中损失立即以幂律衰减下降不同,协同获取是门控的,需要达到配对数据的临界阈值后协同损失才会下降。我们引入了一族新的多模态缩放定律,其中总数据归因损失是四个独立幂律之和,分别对应冗余、每个模态的独特通道和协同四个不同信息通道,并展示了该定律在理论上更合理,并在我们的实验中经验上有效。该定律在我们的实验中比现有定律更准确地预测多模态损失,拟合测试误差为3.2%,而现有定律的最佳配对扩展误差为10.4%。
英文摘要
Existing multimodal scaling laws fit multimodality terms empirically after testing and never vary how much data is multimodally paired at fixed data budgets. We investigate how, under the same total data per modality, changing the number of paired data affects loss curves in multimodal classification tasks. We train models in three different environments and run experiment sweeps varying data sizes and pairing budget. Pairing ratios have a dramatic impact on loss and this impact is directly tied to how much information synergy the task contains. Only paired data is able to reduce synergistic loss, while unpaired data can reduce redundant or unimodal information up until unimodal floors. Unlike traditional scaling laws where loss drops immediately in power law decay, synergy acquisition is gated, requiring a critical threshold of paired data before synergistic loss falls at all. We introduce a new family of multimodal scaling laws where total data-attributable loss is the sum of four individual power laws corresponding to the four different information channels of redundancy, a unique channel per modality, and synergy, and show how this law is both more theoretically sound and empirically valid across our experiments. This law predicts multimodal loss in our experiments more accurately than existing laws, with 3.2% error on fit tests versus 10.4% error for the best pairing extension of published laws.
发表机构
- University of Southern California(南加利福尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。