AI 中文总结
该研究针对规则密集型国家标准文档评审推出首个基准GB/T-Bench,提出多智能体框架GB/T-Reviewer,实验显示其大幅缩小了LLM与专家在该任务上的能力差距,为高风险文档领域可信AI奠定基础。
AI 中文摘要
大语言模型(LLM)越来越多地支持复杂专业任务,但其在规则密集型文档评审方面的能力尚未得到充分评估。中国国家标准(如GB/T标准)是一个具有代表性的测试平台:它们篇幅较长、结构高度规整,且受限于范围、术语、规范措辞及跨章节一致性的明确规则。现有基准侧重于领域知识和问答任务,基本忽略了专业文档的内在质量评审;此类评审高度依赖人类专家,成本高昂且难以规模化。为弥合这一差距,我们推出首个针对国家标准文档结构化评审的基准——GB/T-Bench。其GB/T评审分类法是一个分层架构,涵盖文档结构、范围对齐、规范模态、术语一致性及规范引用,包含25种可诊断的错误类型。一种可控反例生成机制结合确定性规则与受约束的LLM改写,将488份文档处理为7306个可追溯的评审错误实例用于评估。我们还制定了面向诊断的评估协议,要求错误位置、评审维度及错误类型的精确匹配,同时辅以文档级覆盖率指标。我们进一步提出多智能体框架GB/T-Reviewer,该框架将评审知识转化为专业技能,并协调全局检查、针对性诊断、规则扫描及结果验证。对14种主流LLM的实验显示,人类与LLM之间存在显著差距:最强模型的CMCS仅为0.3280,而专家的CMCS为0.6640;GB/T-Reviewer将最佳CMCS提升至0.5094,表明结构化技能协同对规则密集型文档评审具有重要价值。本研究为标准化及其他高风险文档领域的可信AI奠定了基础。
英文摘要
Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit rules for scope, terminology, normative wording, and cross-section consistency. Existing benchmarks focus on domain knowledge and question answering, largely overlooking intrinsic quality review for professional documents. Such reviews rely heavily on human experts, making them costly and difficult to scale. To bridge this gap, we introduce GB/T-Bench, the first benchmark for the structured review of national standard documents. Its GB/T Review Taxonomy is a hierarchical schema covering document structure, scope alignment, normative modality, terminology consistency, and normative references, with 25 diagnosable error types. A controllable counterexample generation mechanism combines deterministic rules and constrained LLM rewriting to process 488 documents into 7,306 traceable review error instances for evaluation. We also develop a diagnosis-oriented evaluation protocol requiring exact matches on error location, review dimension, and error type, plus document-level coverage metrics. We further propose GB/T-Reviewer, a multi-agent framework that converts review knowledge into specialized skills and coordinates global inspection, targeted diagnosis, rule scanning, and result verification. Experiments with 14 mainstream LLMs reveal a substantial human-LLM gap: the strongest model achieves only 0.3280 CMCS versus 0.6640 for experts. GB/T-Reviewer raises the best CMCS to 0.5094, showing the value of structured skill coordination for rule-intensive document review. This work paves the way for trustworthy AI in standardization and other high-stakes document domains.