IndustryLLM:面向工业采购的失败驱动型大语言模型训练
IndustryLLM: Failure-Driven LLM Training for Industrial Procurement
AI总结:
提出IndustryLLM,一种失败驱动的工业采购语言模型,通过CPT和SFT结合多语域改写与证据门控接口,显著提升查询结构化、GMV和满意度并降低延迟。
AI中文摘要:
工业采购要求语言模型在严格的安全容差下,弥合非正式买家行话、稀疏的市场属性和权威工程标准之间的鸿沟。我们提出了IndustryLLM,一个开放权重的工业语言模型,基于Qwen3.5-35B-A3B-Base进行训练(总参数35B,每个token激活约3B参数,视觉编码器冻结)。我们不依赖通用的文本扩展方法,而是引入了一种失败驱动的适配方案,涵盖持续预训练(CPT)和监督微调(SFT)。CPT利用了一个约100B token的精选语料库,其中包含5B token的国家标准(如GB/T)和技术档案、10B token的去标识化真实工业交易和询价记录,以及60B token的通用回放数据。为了克服语域不匹配和事实脆弱性问题,我们通过跨10种体裁和8种写作风格的多语域改写、基于置信度的最小事实编辑以及针对错误的问答合成(解决口语化拼写错误如'42-luo-mu'→42CrMo,扩展模糊代码如'16674'→GB/T 16674,并澄清冲突的尺寸规格),系统性地重建了估计约20B token的领域子集。对于下游部署,我们形式化了一个证据门控的约束评估接口,强制执行三值逻辑,其中未经验证的产品证据保持为“未知”而非“满足”。离线评估显示,在采购查询结构化方面持续改进(精确匹配率提升+2.97个百分点,95%置信区间[2.11, 3.86],在No-Think模式下),而生产环境中的随机在线A/B实验产生了显著改进(+4.25% GMV,+8.3%满意询价),同时延迟从6-7秒降至1.5秒。模型权重和配置已在此https URL发布。
英文摘要:
Industrial procurement requires language models to bridge informal buyer jargon, sparse marketplace attributes, and authoritative engineering standards under strict safety tolerances. We present IndustryLLM, an open-weight industrial language model trained from Qwen3.5-35B-A3B-Base (35B total parameters with ~3B activated per token, with the vision encoder frozen). Rather than relying on generic text scaling, we introduce a failure-driven adaptation recipe spanning continued pre-training (CPT) and supervised fine-tuning (SFT). CPT leverages a curated ~100B-token corpus integrating 5B tokens of national standards (e.g., GB/T) and technical archives, 10B tokens of de-identified real-world industrial transaction and inquiry records, and 60B tokens of general replay. To overcome register mismatch and factual brittleness, we systematically reconstruct an estimated 20B-token domain subset via multi-register rewriting across 10 genres and 8 writing styles, confidence-routed minimal factual editing, and error-targeted QA synthesis (resolving colloquial typos like '42-luo-mu' -> 42CrMo, expanding ambiguous codes like '16674' -> GB/T 16674, and clarifying conflicting dimensional specs). For downstream deployment, we formalize an evidence-gated constraint-evaluation interface enforcing three-valued logic where unverified product evidence remains unknown rather than satisfied. Offline evaluations demonstrate consistent gains on procurement-query structuring (+2.97 percentage points in exact match, 95% CI [2.11, 3.86] in No-Think mode), while randomized online A/B experiments in production yield substantial improvements (+4.25% GMV, +8.3% satisfied inquiries) alongside a latency reduction from 6-7 s to 1.5 s. Model weights and configs are released at https://huggingface.co/alibaba-multimodal-industrial-ai/IndustryLLM.