通过学习和搜索理解组合优化中类人解决方案
Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search
浏览论文内容
中文总结 AI 辅助
研究欧几里得旅行商问题中人类如何找到接近最优解,通过大规模行为和计算研究,对比人类与基于指针网络的神经策略,发现类人解决方案或源于结构化监督学习、强化学习及测试时搜索的结合。
中文摘要 AI 辅助
人类常常能为组合优化问题找到良好解决方案,即便对先进计算机算法而言计算也很困难。在欧几里得旅行商问题(TSP)中,人们能快速给出接近最优的路线。本文通过对欧几里得TSP中人类表现的大规模行为和计算研究来探讨相关问题。采样多种TSP实例,收集人类解决方案并与基于指针网络的神经策略比较。用多种目标训练网络,包括强化学习等。人类路线与最优路线不同但处于接近最优的几何盆地,最佳解释是经最优监督预训练、强化学习微调并通过N选1采样解码的模型。这些发现表明类人解决方案可能源于结构化监督学习、强化学习和测试时搜索的结合。
英文摘要
Humans often find good solutions to combinatorial optimization problems that are computationally hard even for advanced computer algorithms. In the Euclidean traveling salesman problems (TSP), people rapidly produce tours that are near-optimal, despite severe limits on time and computation. What makes a tour human-like, and how might such solutions be learned? Here we address these questions through a large-scale behavioral and computational investigation of human performance in Euclidean TSP. We sampled a broad space of TSP instances, collected human solutions, and compared them with neural policies based on Pointer Networks, which are recurrent neural networks with an attention-based pointing mechanism that define probability distributions over valid tours. We trained these networks under multiple objectives, including reinforcement learning (RL), supervised learning from optimal tours, supervised learning from human tours, and RL fine-tuning after optimal-supervised pretraining. Human tours were not identical to optimal tours, but occupied a near-optimal geometric basin: they shared many structural properties with optimal solutions while preserving systematic human-specific deviations. The best account of human tours was not direct imitation of optimal tours, but a model pretrained on optimal tours, fine-tuned by RL, and decoded through $\text{Best-of-}N$ sampling. These findings suggest that human-like solutions may emerge from a combination of structured supervised learning, RL, and test-time search, echoing computational principles underlying many modern artificial intelligence systems.
发表机构
- The University of Warwick(华威大学)
- The University of Hong Kong(香港大学)
- The Chinese University of Hong Kong(香港中文大学)
- South China Normal University(华南师范大学)
- The University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)
机构由 AI 辅助整理,请以论文原文为准。