Industrial-Instruction:一种从工业技术报告构建指令调优与基准数据集的端到端框架
Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports
- Hamedan University of Technology(哈马丹理工大学)
- University of Antwerp(安特卫普大学)
- Flanders Make Strategic Research Center(弗兰德斯制造战略研究中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对工业技术报告无公开指令调优/基准数据集的问题,提出Industrial-Instruction端到端框架,构建两个QA数据集,对比开放与前沿模型的数据生成效果,为工业领域提供可复现的基准与训练数据路径。
AI中文摘要:
工业技术报告包含用于维护、故障排查和产品工程的高价值知识,但其异构结构(密集散文、规格说明、表格)使得标准检索和问答(QA)管道难以对其进行索引和推理,且目前尚无基于此类文档构建的公开指令调优或基准数据集。我们通过Industrial-Instruction解决这一缺口,贡献:(i)两个从真实工业技术报告构建的开放QA数据集;(ii)生成这些数据集的端到端管道。使用906份公开的松下文档(共7525页),我们应用布局感知提取技术,构建语义检索索引,并基于检索到的证据生成多项选择题QA,涵盖五种查询-文档关系(无关检索、单文档支持、多文档支持、单文档答案、多文档答案)。在过滤初始23900个生成样本后,每个数据集提供约13600个QA对,附带源文档和保留的基准拆分。对规模在100亿参数以下的小型开放大语言模型(LLM)进行微调后,在松下基准上,集匹配准确率从28.5%提升至42.0%,F1值从46.6%提升至63.5%。我们通过同一管道发布两个并行版本:一个由开放权重模型Qwen3-30B-A3B-Instruct生成,另一个由基于API的闭源模型Claude-Opus-4.6生成,可直接对比开放模型与前沿模型的数据生成效果。Claude-Opus-4.6生成的数据集产出更干净的原始语料库,微调增益更大,但成本约高出两个数量级。MMLU评估显示,在Claude-Opus-4.6数据集上训练的模型基本保留所有通用知识,而在Qwen生成的数据集上训练的模型则出现少量但可测量的遗忘效应。综上,这些数据集和管道为从真实文档构建可扩展工业基准和训练数据提供了实用、可复现的路径。
英文摘要:
Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, but their heterogeneous structure (dense prose, specifications, tables) makes them difficult to index and reason over with standard retrieval and QA pipelines, and no public instruction-tuning or benchmark datasets are built from such documents. We address this gap with Industrial-Instruction, contributing (i) two open QA datasets built from real industrial technical reports and (ii) the end-to-end pipeline that produces them. Using 906 public Panasonic documents (7,525 pages), we apply layout-aware extraction, build a semantic retrieval index, and synthesize multiple-choice QA grounded in retrieved evidence under five query-document relationships (irrelevant retrieval, single-/multi-document support, single-/multi-document answer). After filtering an initial 23.9k generated samples, each dataset provides approximately 13.6k QA pairs with source documents and a held-out benchmark split. Fine-tuning small open LLMs (under 10B parameters) improves Set-Match Accuracy from 28.5% to 42.0% and F1 from 46.6% to 63.5% on the Panasonic benchmark. We release two parallel versions built by the same pipeline: one generated with the open-weight Qwen3-30B-A3B-Instruct model and one with the closed, API-based Claude-Opus-4.6 model, enabling a direct comparison of open- versus frontier-model data generation. The Claude-Opus-4.6 dataset yields a cleaner raw corpus and larger fine-tuning gains, at roughly two orders of magnitude higher cost. MMLU evaluation shows models trained on the Claude-Opus-4.6 data retain essentially all general knowledge, versus a small but measurable forgetting effect for the Qwen-generated data. Together, these datasets and pipeline offer a practical, reproducible path toward scalable industrial benchmarks and training data from real-world documentation.