发表机构
Soochow University; School of Future Science and Engineering(苏州大学; 未来科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多智能体组合优化的性能局限,提出GeoPAR框架,通过三个关键组件优化几何建模与冲突处理,在两类问题上实现了更好的大规模泛化与高效推理。
AI 中文摘要
多智能体组合优化问题因具有NP难特性而极具挑战性。近期的并行自回归神经求解器通过允许智能体同时决策提升了推理效率,但在大规模实例上的性能往往会下降,这主要归因于对局部几何结构的建模不足,以及仅在动作生成后才处理冲突的任务选择。为解决这些局限,我们提出了GeoPAR,这是一种用于可扩展多智能体组合优化的几何引导并行自回归强化学习框架。GeoPAR集成了三个关键组件:(1)投影窗口稀疏几何机制,通过多方向投影构建轻量型局部候选邻域;(2)稀疏边偏置注意力,将这些几何关系注入节点表示;(3)缓存引导的冲突感知分配,在解码过程中复用几何缓存以抑制互斥任务的重复选择。在异构车辆路径规划和开放多车场取送问题上的实验表明,GeoPAR提升了大规模零样本泛化能力,同时大幅减少了回退步数并保持了高效推理。
英文摘要
Multi-agent combinatorial optimization problems are notoriously challenging due to their NP-hard nature. Recent parallel autoregressive neural solvers improve inference efficiency by allowing agents to make decisions simultaneously, but their performance often degrades on large-scale instances. This is largely attributable to weak modeling of local geometric structures and the fact that conflicting task selections are handled only after action generation. To address these limitations, we propose GeoPAR, a geometry-guided parallel autoregressive reinforcement learning framework for scalable multi-agent combinatorial optimization. GeoPAR integrates three key components: (1) a projection-window sparse geometry mechanism that builds lightweight local candidate neighborhoods through multi-directional projections, (2) sparse edge-biased attention that injects these geometric relations into node representations, and (3) cache-guided conflict-aware assignment that reuses the geometric cache during decoding to suppress duplicate selections of exclusive tasks. Experiments on heterogeneous vehicle routing and open multi-depot pickup-and-delivery problems show that GeoPAR improves large-scale zero-shot generalization while substantially reducing rollout steps and maintaining efficient inference.