arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19761math.OC

通过最优传输利用异构数据进行条件优化

Harnessing Heterogeneous Data for Conditional Optimization via Optimal Transport

Jonathan Yu-Meng Li, Qinyu Wu

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对条件优化中从有限样本学习条件分布困难的问题,开发基于最优传输的分布鲁棒框架,提出三个OT模糊集,推导相关表述、条件等,通过条件分类问题证明框架价值。

中文摘要 AI 辅助

条件优化根据上下文或事件信息调整决策,但其实际应用常受限于从目标联合分布的有限样本中学习相关条件分布的困难。当目标联合数据稀缺或不可用时,或者很少有观测值落在感兴趣的条件区域时,这一挑战尤为严峻。相关联合数据可能来自多个来源,但这些来源可能相对于目标分布存在偏差,不能简单合并。我们开发了一个基于最优传输(OT)的分布鲁棒框架,用于在条件优化中利用此类异构数据。该框架使用OT距离到经验源分布来构建联合分布上的模糊集,并在合理的目标定律上优化最坏情况条件性能。我们提出了三个OT模糊集,它们捕获了使用异构源的不同方式:强制同时源一致性、通过权重聚合源差异以及将模糊集以OT重心为中心。我们推导了易于处理的重新表述,建立了可行性条件,讨论了参数选择,并刻画了这些表述之间的关系,揭示了鲁棒性、信息聚合和计算复杂性之间的权衡。我们通过使用来自多个商店的需求和产品特征数据的条件分类问题证明了该框架的价值。

英文摘要

Conditional optimization tailors decisions to contextual or event information, but its practical use is often limited by the difficulty of learning the relevant conditional distribution from finite samples of a target joint distribution. This challenge is especially acute when target joint data are scarce or unavailable, or when few observations fall in the conditioning region of interest. Related joint data may be available from multiple sources, such as different stores, markets, populations, or operating environments, but these sources may be biased relative to the target distribution and cannot be pooled naively. We develop a distributionally robust framework based on optimal transport (OT) for harnessing such heterogeneous data in conditional optimization. The framework constructs ambiguity sets over joint distributions using OT distances to empirical source distributions and optimizes worst-case conditional performance over plausible target laws. We propose three OT ambiguity sets that capture different ways of using heterogeneous sources: enforcing simultaneous source consistency, aggregating source discrepancies through weights, and centering the ambiguity set at an OT barycenter. We derive tractable reformulations, establish feasibility conditions, discuss parameter choices, and characterize the relationships among the formulations, revealing trade-offs between robustness, information aggregation, and computational complexity. We demonstrate the value of the framework through a conditional assortment problem using demand and product-feature data from multiple stores.

↑