arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31871cs.AIcs.CL

IndustryLLM:面向工业采购的失败驱动型大语言模型训练

IndustryLLM: Failure-Driven LLM Training for Industrial Procurement

Liang Ding, Zhiang Xu, Yuyang Sheng, Bin Chen, Songlin Bai, Run Zhu, Dingjun Wu, Hui Xu, Yandi Wang, Fulin Shi, Leilei Gan, Linlin Yu, Qihuang Zhong, Keqin Peng… 展开作者

Liang Ding, Zhiang Xu, Yuyang Sheng, Bin Chen, Songlin Bai, Run Zhu, Dingjun Wu, Hui Xu, Yandi Wang, Fulin Shi, Leilei Gan, Linlin Yu, Qihuang Zhong, Keqin Peng, Yalong Li, Chengfu Huo

AI总结:

提出IndustryLLM,一种失败驱动的工业采购语言模型,通过CPT和SFT结合多语域改写与证据门控接口,显著提升查询结构化、GMV和满意度并降低延迟。

AI中文摘要:

工业采购要求语言模型在严格的安全容差下,弥合非正式买家行话、稀疏的市场属性和权威工程标准之间的鸿沟。我们提出了IndustryLLM,一个开放权重的工业语言模型,基于Qwen3.5-35B-A3B-Base进行训练(总参数35B,每个token激活约3B参数,视觉编码器冻结)。我们不依赖通用的文本扩展方法,而是引入了一种失败驱动的适配方案,涵盖持续预训练(CPT)和监督微调(SFT)。CPT利用了一个约100B token的精选语料库,其中包含5B token的国家标准(如GB/T)和技术档案、10B token的去标识化真实工业交易和询价记录,以及60B token的通用回放数据。为了克服语域不匹配和事实脆弱性问题,我们通过跨10种体裁和8种写作风格的多语域改写、基于置信度的最小事实编辑以及针对错误的问答合成(解决口语化拼写错误如'42-luo-mu'→42CrMo,扩展模糊代码如'16674'→GB/T 16674,并澄清冲突的尺寸规格),系统性地重建了估计约20B token的领域子集。对于下游部署,我们形式化了一个证据门控的约束评估接口,强制执行三值逻辑,其中未经验证的产品证据保持为“未知”而非“满足”。离线评估显示,在采购查询结构化方面持续改进(精确匹配率提升+2.97个百分点,95%置信区间[2.11, 3.86],在No-Think模式下),而生产环境中的随机在线A/B实验产生了显著改进(+4.25% GMV,+8.3%满意询价),同时延迟从6-7秒降至1.5秒。模型权重和配置已在此https URL发布。

英文摘要:

Industrial procurement requires language models to bridge informal buyer jargon, sparse marketplace attributes, and authoritative engineering standards under strict safety tolerances. We present IndustryLLM, an open-weight industrial language model trained from Qwen3.5-35B-A3B-Base (35B total parameters with ~3B activated per token, with the vision encoder frozen). Rather than relying on generic text scaling, we introduce a failure-driven adaptation recipe spanning continued pre-training (CPT) and supervised fine-tuning (SFT). CPT leverages a curated ~100B-token corpus integrating 5B tokens of national standards (e.g., GB/T) and technical archives, 10B tokens of de-identified real-world industrial transaction and inquiry records, and 60B tokens of general replay. To overcome register mismatch and factual brittleness, we systematically reconstruct an estimated 20B-token domain subset via multi-register rewriting across 10 genres and 8 writing styles, confidence-routed minimal factual editing, and error-targeted QA synthesis (resolving colloquial typos like '42-luo-mu' -> 42CrMo, expanding ambiguous codes like '16674' -> GB/T 16674, and clarifying conflicting dimensional specs). For downstream deployment, we formalize an evidence-gated constraint-evaluation interface enforcing three-valued logic where unverified product evidence remains unknown rather than satisfied. Offline evaluations demonstrate consistent gains on procurement-query structuring (+2.97 percentage points in exact match, 95% CI [2.11, 3.86] in No-Think mode), while randomized online A/B experiments in production yield substantial improvements (+4.25% GMV, +8.3% satisfied inquiries) alongside a latency reduction from 6-7 s to 1.5 s. Model weights and configs are released at https://huggingface.co/alibaba-multimodal-industrial-ai/IndustryLLM.

补充信息

↑