发表机构
Continker(Continker)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MetroLLM-Bench是一个955案例的基准,评估语言模型作为交通信息亭策略层,涵盖六个地铁系统及十一类任务,通过双层评分(确定性加语义)测试,发现4B PEFT学生模型在Tier 1上超越GPT-5.6,且开源了基准和微调模型。
AI 中文摘要
我们推出了MetroLLM-Bench,一个包含955个案例的基准测试,用于测试语言模型作为交通信息亭策略层的性能。该基准覆盖了六个真实地铁系统,车站数量从37个到414个不等,并涵盖十一个类别,包括路线规划、票价计算、中断处理、无障碍设施和对抗性输入。在每个案例中,模型必须调用结构化工具,并提交一个机器可渲染的终端状态,其中包含结果、适用时的每张票的票价报价以及信息亭操作。十四个确定性评分组件构成第一层;八个语义质量组件构成第二层,其中六个使用语言模型评判器。我们报告第一层得分以及两层的综合得分。分层75/25划分保留了717个案例用于训练数据生成,238个案例用于留出评估。我们评估了来自六个供应商的二十六个模型,其中二十三个被排名。在留出分区上,通过参数高效微调(PEFT)训练的4B Qwen 3.5学生模型在第一层上超过了GPT-5.6的两个层级(91.3对90.6和90.0),并在最大推理努力下与GPT-5.4完整版持平(91.4),其Q4_K_M占用空间为2.6 GB。在此训练规模下,更大的9B和27B学生模型相对于4B学生模型没有提供进一步的第一层改进。在四个Qwen规模中,PEFT相对于相应基础模型的增益从2B时的+7.03分(三个训练种子)下降到27B时的-0.91分;每个种子在每个规模上都显示出相同的方向。一个确定性的基于规则的基线在第一层达到84.6,剩余的语言模型优势集中在策略适应、复合场景、无障碍设施和时间推理方面。Muse Glimmer 30B在综合排名中领先,仅服务配置就使Qwen 3.5到3.8的比较在第一层上移动了2.7分。该基准、测试工具、复现指南和微调后的学生模型已在此https URL发布。
英文摘要
We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a language-model judge. We report Tier 1 and the combined score of both tiers. A stratified 75/25 split reserves 717 cases for training-data generation and 238 for held-out evaluation. We evaluate twenty-six models from six vendors, of which twenty-three are ranked. On the held-out partition, a 4B Qwen 3.5 student trained through parameter-efficient fine-tuning (PEFT) exceeds both GPT-5.6 tiers on Tier 1 (91.3 against 90.6 and 90.0) and matches GPT-5.4 full at maximum reasoning effort (91.4), with a 2.6 GB Q4_K_M footprint. Larger 9B and 27B students provide no further Tier 1 improvement over the 4B student at this training scale. Across the four Qwen sizes, the PEFT gain over the corresponding base model decreases from +7.03 points at 2B (three training seeds) to -0.91 at 27B; every seed shows the same direction at every size. A deterministic rule-based baseline reaches 84.6 on Tier 1, with the remaining language-model advantage concentrated in policy adaptation, compound scenarios, accessibility, and temporal reasoning. Muse Glimmer 30B leads the composite ranking, and serving configuration alone moves the Qwen 3.5-to-3.8 comparison by 2.7 Tier 1 points. The benchmark, harness, reproduction guide, and fine-tuned students are released at https://github.com/continker/metrollm-bench.
Comments23 pages, 5 figures, 10 tables. Code and data at https://github.com/continker/metrollm-bench (tag paper-v1.2); DOI 10.5281/zenodo.21893944