arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

决策中的弱到强学习

Weak-to-Strong Learning in Decision Making

Jingwei Ji, Renyuan Xu

arXiv 2607.18467首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对运营决策中预测模型训练的数据不对称问题,提出决策感知弱到强(W2S)框架,利用有标签和无标签数据训练模型,建立相关风险上下界得出性能改善条件,通过实验验证理论。

AI 中文摘要

许多运营决策依赖于根据可观察情境估计不确定结果的预测模型。然而,训练此类模型常面临数据不对称问题:有标签结果稀缺或获取成本高,而情境协变量丰富。基于此,我们开发了决策感知弱到强(W2S)框架,利用有标签和无标签数据改进情境随机优化。具体先利用有限有标签数据训练弱模型,再用其在无标签情境上生成预测结果分布,为训练强模型提供软监督。我们建立了W2S超额决策风险的非渐近上界和仅强模型基准的互补下界,比较得出W2S改善下游决策性能的明确充分条件。关键量是弱特征与强特征表示之间的相关维度:当它小时,丰富的无标签数据可减少教师错误在非重叠方向上的影响。合成报童实验和基于真实数据的评论审核实验提供了与理论一致的经验证据。

英文摘要

Many operational decisions rely on predictive models that estimate uncertain outcomes conditional on observable contexts. Training such models, however, often faces a fundamental data asymmetry: labeled outcomes are scarce or costly to obtain, while contextual covariates are abundant. Motivated by this data asymmetry, we develop a decision-aware weak-to-strong (W2S) framework that leverages both labeled and unlabeled data to improve contextual stochastic optimization. Specifically, we first train a weak model using limited labeled data and then use it to generate predicted outcome distributions on unlabeled contexts. These distributions provide soft supervision for training a strong model. We establish a non-asymptotic upper bound on the excess decision risk of W2S and a complementary lower bound for a strong-only benchmark. Their comparison yields explicit sufficient conditions under which W2S improves downstream decision performance. The key quantity is the correlation dimension between the weak and strong feature representations: when it is small, abundant unlabeled data reduce the effect of teacher errors along non-overlapping directions. A synthetic newsvendor experiment and a comment moderation experiment based on real-world data provide empirical evidence consistent with the theory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑