发表机构
NCB Hazcheck Limited; Durham University(NCB Hazcheck有限公司; 杜伦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建首个IMDG规则合规性基准DGEval,评估13个LLM的性能,发现其在积载等安全关键领域表现弱,仅可辅助结构化查询,需人工监督。
AI 中文摘要
海运危险货物运输是受《国际海运危险货物规则》(IMDG Code)管辖的高后果活动,该规则是一个复杂的监管框架,分类、包装、积载或隔离方面的错误可能导致火灾、爆炸、有毒物质释放,或造成人员伤亡或船舶损失。正确的合规需要准确解读数百页相互关联的条款,这些条款每两年更新一次修正案。从业者越来越多地使用大型语言模型(LLMs)作为决策支持工具,但目前尚无系统评估验证它们是否能可靠解读IMDG规则以用于安全关键用途。本文介绍DGEval,这是首个用于评估LLMs对IMDG 42-24修正案知识的基准。该基准基于NCB Hazcheck在线学习平台上专家编写的问题以及危险货物清单(DGL)的结构化查询构建,包含1678个问题,涵盖多项选择题、开放式问题、DGL查询和监管识别任务。我们评估了来自六个提供商的13个模型,涉及多种推理配置,包括一个海运领域特定微调模型,并测试了网络搜索的效果。尽管表现最佳的模型在多项选择题上超过了人类从业者基线,但所有模型在操作上安全关键的积载、隔离和监管召回领域表现最弱。这些结果表明,LLMs可支持合规任务,尤其是结合网络搜索的结构化DGL查询,但在操作领域和监管文本召回方面的不可靠性意味着,在安全关键环境中部署前,仍需人工监督和权威来源验证。DGEval被设计为一种安全保障工具,可随着模型发展持续应用,而非对当前能力的最终定性。
英文摘要
The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code, a complex regulatory framework where errors in classification, packaging, stowage, or segregation can result in fire, explosion, toxic release, or loss of life or vessel. Correct compliance requires accurately interpreting hundreds of pages of interacting provisions, updated on a two-year amendment cycle. Practitioners increasingly use Large Language Models (LLMs) as decision-support tools, yet no systematic evaluation exists of whether they can reliably interpret IMDG requirements for safety-critical use. This paper introduces DGEval, the first benchmark for evaluating LLM knowledge of IMDG Amendment 42-24. Built from expert-written questions on a commercial e-learning platform and structured lookups from the Dangerous Goods List (DGL), it comprises 1,678 questions across multiple-choice, open-ended, DGL lookup, and regulatory identification tasks. We evaluate 13 models from six providers across multiple thinking configurations, including one maritime domain-specific fine-tuned model, and test the effect of web search. Although the best-performing model exceeds the human practitioner baseline on multiple-choice questions, all models are weakest in the operationally safety-critical areas of stowage, segregation, and regulatory recall. These results indicate that LLMs may support compliance tasks, particularly structured DGL lookups with web search, but unreliability in operational areas and regulatory-text recall means human oversight and authoritative source verification remain necessary before deployment in any safety-critical context. DGEval is designed as a safety assurance instrument to be applied continuously as models evolve, not as a settled characterisation of current capability.
Comments28 pages, 2 figures