arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26375cs.LGcs.AIcs.DM

CG4AI:一种用于在约束条件下训练AI模型的列生成框架

CG4AI: A Column Generation Framework for Training AI Models Under Constraints

  • Huawei Technologies Ltd., France Research Center(华为技术有限公司法国研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Youcef Magnouche, Abderrahmane Driouch, Sébastien Martin, Pierre Bauguion

中文总结 AI 辅助

CG4AI是带线性约束的AI模型训练的列生成框架,通过主LP、定价子问题和割平面程序实现,在MNIST和SNDLIB网络上验证,可生成可行预测器且准确率优于单模型基线。

中文摘要 AI 辅助

标准机器学习训练会在数据集上最小化损失函数,但无法保证得到的模型满足对其输出的预定义规则或约束。在从自主系统到网络路由的许多实际应用中,此类保证至关重要。我们提出CG4AI,这一框架在对组合输出施加线性约束的同时,构建AI模型的凸组合。主线性规划(LP)确定最优混合权重,而定价子问题在LP对偶变量的引导下生成新模型,重点关注违反最严重的约束。割平面程序将可行性保证扩展到训练集之外。我们将CG4AI应用于两个问题:(i)MNIST上的数字分类,我们展示了约束的四种不同用途:仅从约束中学习、提高对抗鲁棒性、纠正误分类示例以及强制输出重标记;(ii)多商品流问题,在该问题中,对神经网络路由预测器施加链路容量约束。在MNIST和标准SNDLIB基准网络上的实验表明,CG4AI可可靠生成可行的预测器,同时比单模型基线实现更好的准确率。

英文摘要

Standard machine-learning training minimizes a loss function over a dataset, but does not guarantee that the resulting model will satisfy predefined rules or constraints on its outputs. In many real-world applications, ranging from autonomous systems to network routing, such guarantees are essential. We propose CG4AI, a framework that builds a convex combination of AI models while enforcing linear constraints on the combined output. A master linear program (LP) determines the optimal mixture weights, while a pricing subproblem generates new models guided by LP dual variables, focusing attention on the most violated constraints. A cutting-plane procedure extends feasibility guarantees beyond the training set. We apply CG4AI to two problems: (i) digit classification on MNIST, where we demonstrate four distinct uses of constraints, learning from constraints alone, improving adversarial robustness, correcting misclassified examples, and enforcing output relabeling; and (ii) the multi-commodity flow problem, where link capacity constraints are enforced on neural-network routing predictors. Experiments on MNIST and standard SNDLIB benchmark networks show that CG4AI reliably produces feasible predictors while achieving better accuracy than single-model baselines.

↑