arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于图神经网络学习最优动态匹配

Learning Optimal Dynamic Matching via Graph Neural Networks

Genta Okada, Shunya Noda, Junpei Komiyama, Akira Matsushita

arXiv 2607.28925首次发表:更新:

AI 中文总结

该研究针对动态匹配市场问题,开发基于残差图的图神经网络强化学习框架,在两类基准中均优于传统贪心匹配策略,可自适应调整匹配决策。

AI 中文摘要

动态匹配市场需要做出匹配对象与时机的决策:当前匹配会产生价值,但会移除可能创造更好未来机会的参与者。我们针对有限、演化的加权图,开发了一种基于价值的强化学习框架来解决该问题。我们研究了具有随机到达、节点类型转换、边实现和外生退出的无限时域连续时间模型。我们证明了事件时间约简:不失最优性,规划者在每个外生事件后立即行动,然后等待下一个事件。我们进一步表明,最优的逐边Q函数由决策后残差图上的单个延续值函数表征,将学习对象从状态-动作值约简为图值。精确的动作选择仍需要组合匹配优化;我们用图神经网络近似该值,通过时间差分学习训练它,并将其用于前向贪心匹配启发式算法。在二元类型基准中,所学策略通过为有价值匹配的稀有到达保留常见节点,仅在密集池形成较低价值匹配,显著优于即时和阈值贪心规则。在肾脏配对捐赠基准中,当退出不可预测时,其表现与即时贪心相似;当警告可靠时,它恢复了患者匹配的逻辑;在中等警告概率下,它优于即时贪心和患者贪心中较好的那个。这些结果表明,残差图值学习产生了依赖于状态的动态匹配策略,可适应已实现的连通性和退出信息。

英文摘要

Dynamic matching markets require decisions about whom to match and when: matching now yields value but removes participants who may create better future opportunities. We develop a value-based reinforcement-learning framework for this problem on finite, evolving weighted graphs. We study an infinite-horizon continuous-time model with stochastic arrivals, node-type transitions, edge realizations, and exogenous exits. We prove an event-time reduction: without loss of optimality, the planner acts immediately after each exogenous event and then waits for the next one. We further show that the optimal edge-wise $Q$-function is characterized by a single continuation-value function on post-decision residual graphs, reducing the learned object from state-action values to graph values. Exact action selection still requires combinatorial matching optimization; we approximate the value with a graph neural network, train it by temporal-difference learning, and use it in a forward-greedy matching heuristic. In a binary-type benchmark, the learned policy substantially outperforms immediate and threshold-greedy rules by preserving common nodes for rare arrivals of valuable matches while forming lower-value matches only in thick pools. In a kidney paired donation benchmark, it performs similarly to immediate greedy when exits are unpredictable, recovers the logic of patient matching when warnings are reliable, and outperforms the better of Immediate Greedy and Patient Greedy across intermediate warning probabilities. These results show that residual-graph value learning yields state-dependent dynamic matching policies that adapt to realized connectivity and exit information.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑