arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于模拟轨迹的大语言模型引导式启发式设计:动态生产与自动导引车调度的案例研究

LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling

Jinbo Li, Chuanhao Li

arXiv 2608.09343首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出LLM引导的启发式设计框架,通过模拟轨迹诊断优化AGV与生产调度策略,在多组实验中优于多种基线方法,证实模拟轨迹可指导代码层面的策略改进。

AI 中文摘要

基于模拟的优化(SBO)在随机动态环境下评估可执行策略,但多数方法将模拟器视为黑箱:仅通过聚合分数对候选策略排序,不揭示其失败原因或需调整的策略逻辑。我们提出一种大语言模型(LLM)引导式启发式设计框架,该框架利用重复模拟进行选择,并借助事件级轨迹开展诊断。每个当前最优策略通过多次重复评估,重放其得分最低的一次重复会生成可查询的轨迹;一个管理智能体从该轨迹证据中构建瓶颈假设,编辑智能体则在代码层面实施并行修订。经执行检查和重复评估后,采用当前最优选择机制仅保留改进项。LLM修订在评估批次间进行,每次模拟运行由固定策略控制。我们在动态生产与自动导引车(AGV)调度的离散事件模拟中评估该框架:使用Gemini-3.1-Pro开展5次独立优化运行,最终平均得分在模拟器0-100分制下平均为77.51;在得分最高的运行中,基于轨迹的诊断推动了主动充电、距离感知AGV分配及重新平衡的调度优先级,使当前最优平均得分从62.49提升至78.61。在100个匹配种子上,最优最终策略在所有种子上均优于代表性的滚动混合整数线性规划(rolling-MILP)、基于规则及元启发式策略,且在随机故障下无需重新优化即可保持优势;在针对更长时间范围和可变到达间隔时间分别重新优化后,所得策略再次优于所有基线。对两种LLM主干的 ablation 实验显示,移除并行候选生成或轨迹数据库访问任一环节均会降低最终平均得分。这些结果表明,模拟轨迹可指导复杂基于模拟的调度场景中针对性的代码层面策略改进。

英文摘要

Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most methods treat the simulator as a black box: aggregate scores rank candidates without revealing why they fail or which policy logic should change. We present an LLM-guided heuristic design framework that uses repeated simulation for selection and event-level traces for diagnosis. Each incumbent is assessed through multiple replications, while replaying its lowest-scoring one produces a queryable trace. A manager agent formulates bottleneck hypotheses from this evidence, and editing agents implement parallel code-level revisions. After execution checks and repeated evaluation, best-so-far selection retains only improvements. LLM revision occurs between evaluation batches, while a fixed policy controls each simulation run. We evaluate the framework in a discrete-event simulation of dynamic production and automated guided vehicle (AGV) scheduling. Across five independent optimization runs with Gemini-3.1-Pro, final mean scores averaged 77.51 on the simulator's 0-100 scale. In the highest-scoring run, trace-based diagnoses motivated proactive charging, distance-aware AGV assignment, and rebalanced dispatch priorities, raising the best-so-far mean score from 62.49 to 78.61. On 100 matched seeds, the best final policy outscored representative rolling-MILP, rule-based, and metaheuristic policies on every seed and retained its advantage under random faults without re-optimization. After separate re-optimization for a longer horizon and variable order interarrival times, the resulting policies again outscored all baselines. Ablations with two LLM backbones showed that removing either parallel candidate generation or trace-database access reduced final mean scores. These results show that simulation traces can guide targeted code-level policy improvement in complex simulation-based scheduling.

Comments33 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑