发表机构
University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SCHEDBench基准针对13个LLMs的测试表明,模型对同一调度问题的语义等价表述缺乏不变性,约束重排序引发的敏感性最为突出。
AI 中文摘要
本文介绍SCHEDBench,这是一个用于评估表面形式变化下组合调度约束忠实度的自然语言基准。SCHEDBench基于规范调度实例和求解器生成的可行性与最优性,用于评估大语言模型(LLMs)在不同自然语言(NL)表面形式下生成的调度是否具有相同的约束可行行为。SCHEDBench涵盖1132个实例,涉及作业车间调度问题(JSP)、单模式与多模式资源受限项目调度问题(RCPSP)、护士排班/调度以及不同难度的课程排课问题。这些实例通过领域特定模板、主题实体、词汇-句法模板改写以及约束级表面形式变化被模板化为自然语言问题,参考解已验证可行性和目标最优性。对13个前沿及开放权重LLMs的测试发现,模型对同一调度问题的语义等价表述不具备可靠不变性,表面形式变化会降低可行性,并在匹配实例上引发高于噪声的单实例硬约束违反偏移;在 tested 的孤立维度中,约束重排序产生的高于噪声敏感性最为明显。
英文摘要
This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form variation. Grounded in canonical scheduling instances and solver-derived feasibility and optimality, SCHEDBench assesses whether large language models (LLMs) generate schedules with the same constraint-feasible behavior across varied natural-language (NL) surface forms. SCHEDBench spans 1,132 instances across job-shop scheduling problems (JSP), single and multi-mode resource-constrained project scheduling problems (RCPSP), nurse rostering/scheduling, and curriculum timetabling problems of varying difficulty. Instances are templated into natural language problems using domain-specific templates, themed entities, lexical-syntactic template rephrasing, and constraint-level surface-form variation, with reference solutions verified for feasibility and objective optimality. Across thirteen frontier and open-weight LLMs, we find that models are not reliably invariant to semantically equivalent renderings of the same scheduling problem. Surface-form variation reduces feasibility and induces above-noise shifts in per-instance hard-constraint violations on matched instances. Among the tested isolated axes, constraint reordering yields the clearest above-noise sensitivity.
CommentsEMNLP 2026 Findings