AI 中文总结
本文提出一种通用的视觉-语言求解器,通过为组合优化任务添加视觉表示并采用监督微调和强化学习训练,显著提升了解质量,尤其在复杂和大规模问题上表现突出。
AI 中文摘要
大型语言模型(LLMs)为端到端组合优化(CO)提供了统一接口,但仅靠文本序列化可能会掩盖对生成有效CO解决方案至关重要的空间和关系结构。本文提出了一种通用的视觉-语言求解器,该求解器用输入派生的视觉表示来增强文本实例描述。一个单一的视觉语言模型(VLM)被应用于不同的CO任务,并通过监督微调及后续的验证器引导强化学习进行训练。虽然视觉输入不包含黄金解或解派生信息,但我们的实验表明,与纯文本对应模型相比,VLM通常能提高解决方案质量,在更复杂的CO问题(如CVRP和JSSP)上尤其有显著提升。视觉信息的优势在大规模问题上更为明显。
英文摘要
Large language models (LLMs) have provided a unified interface for end-to-end combinatorial optimization (CO), but textual serialization alone may obscure spatial and relational structures that are important for generating effective CO solutions. This paper presents a general-purpose vision-language solver that augments textual instance descriptions with input-derived visual representations. A single vision-language model (VLM) is applied across different CO tasks and trained using supervised fine-tuning followed by verifier-guided reinforcement learning. While the visual inputs contain no gold solutions or solution-derived information, our experiments show that the VLM generally improves solution quality over its text-only counterpart, with particularly clear gains on more complex CO problems such as CVRP and JSSP. The advantage of visual information is more pronounced at large problem scales.