发表机构
School of Artificial Intelligence, Jilin University; Key Laboratory of Symbolic Computation and Knowledge Engineering, Jilin University(吉林大学人工智能学院; 吉林大学符号计算与知识工程重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CALIBUDGET提出一种带下限保护、可靠性感知的整数分配器,通过校准分割结合模型需求与稳定性,在固定预算下优化异构数据源配额,显著提升混合推理适配性能。
AI 中文摘要
从异构数据源进行固定预算适配时,不仅需要决定使用多少数据,还需要决定每个数据源获得多少曝光量。按规模比例分配的规则可能会排挤小数据源,而仅基于难度的规则可能会追逐噪声估计或将剩余预算分配给接近饱和的数据池。我们引入了CALIBUDGET,一种具有下限保护、可靠性感知的整数分配器,将数据源曝光量视为一个显式的适配变量。通过从训练集内部划分的小型校准子集,它结合了模型需求、下限后的可用性和自助法稳定性,在不改变模型、目标或总预算的情况下,生成精确的、满足容量限制的配额。在一个结合数学和常识数据的受控环境中,CALIBUDGET在所有三次配对的LLaMA-2-7B LoRA+运行中,相较于最强的匹配对照方法(带下限的验证误差),在CommonAvg、FragileAvg和MacroAvg上均有所提升。相应的平均增益分别为0.56、0.46和0.41个百分点(pp)。总体平均增益增加0.18个百分点,而MathAvg下降0.20个百分点,这揭示了覆盖与保留之间的边界,而非均匀的增益。CALIBUDGET仅改变了1.14-1.42%的数据源预算,但在常识任务和种子的24项比较中,有15项性能得到提升。这些结果表明,数据源配额的微小变化可能产生重要影响;而示例级选择则可以决定哪些示例填充每个配额。
英文摘要
Fixed-budget adaptation from heterogeneous data sources requires deciding not only how much data to use, but how much exposure each source receives. Size-proportional rules can crowd out small sources, whereas difficulty-only rules can chase noisy estimates or allocate residual budget to nearly saturated pools. We introduce CALIBUDGET, a floor-protected, reliability-aware integer allocator that treats source exposure as an explicit adaptation variable. From small train-internal calibration splits, it combines model need, post-floor availability, and bootstrap stability, then produces exact capacity-respecting quotas without changing the model, objective, or total budget. In a controlled setting combining mathematical and commonsense data, CALIBUDGET improves CommonAvg, FragileAvg, and MacroAvg over validation-error-with-floor, the strongest matched comparator, in all three paired LLaMA-2-7B LoRA+ runs. The respective mean gains are 0.56, 0.46, and 0.41 percentage points (pp). Overall increases by 0.18 pp, whereas MathAvg decreases by 0.20 pp, exposing a coverage-retention boundary rather than a uniform gain. CALIBUDGET changes only 1.14-1.42% of the source budget but improves performance in 15 of 24 comparisons across commonsense tasks and seeds. These results suggest that small changes in source quotas can matter; example-level selection can then determine which examples fill each quota.