发表机构
Aalborg University(奥尔堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究对比三种适配Qwen2.5-32B-Instruct模型的方法,在污水处理因果问答任务中,DRR检索器兼具高准确率、跨工厂迁移能力及训练高效性,优于其他方法与基线。
AI 中文摘要
污水处理操作员在提出“为何N₂O浓度上升?”或“若将曝气量削减20%会产生什么影响?”等因果问题时,需要基于自身工厂变量的相互作用方式及效应传播速度的答案,而非通用预训练文本。我们对比了三种将冻结的Qwen2.5-32B-Instruct模型适配至具备架构可解释性的污水处理模拟器CCSS-IX的具体方法:实时模拟器专家(方法1)、结构化参数注入(方法2)以及解耦式检索-推理(DRR)检索器(方法3)。在含198个问题的因果基准测试中,三种方法的准确率分别达到99.5%、79%和75.8%,形成了优于最强检索增强基线(48%)的部署阶梯。DRR检索器拥有1.1亿参数,针对单个工厂的训练耗时约17秒;在跨工厂迁移至生物特性不同的工厂后,其准确率仍可达88%,而方法2的静态表格无法实现迁移。在含60个问题的反事实基准测试中,仅方法3能处理干预后的查询:较方法2高出16.3个百分点,配对95%置信区间为[+7.1, +26.4]个百分点,且在时间尺度和运行工况类别上达到100%准确率。在采用OpenBookQA事实语料的AI2推理挑战(ARC)中,相同的选择性检索机制达到79%准确率,而无约束的Llama-3.1-8B为76%、全注入方法为74%,这一+3个百分点的域外复现结果表明该效果并非污水处理领域特有。我们提供了首个针对工业因果问答的实时工具使用、静态参数注入及学习型数值参数检索的单模拟器对比研究。
英文摘要
Wastewater operators need answers grounded in how their plant's variables interact and how fast effects propagate, not in generic pretraining text, when asking causal questions such as "why is N2O rising?" or "what happens if I cut aeration by 20%?". We compare three concrete ways to ground a frozen Qwen2.5-32B-Instruct model in an architecturally interpretable wastewater simulator (CCSS-IX): a live simulator oracle (Method 1), structured parameter injection (Method 2), and a Decoupled Recall-Reasoning (DRR) retriever (Method 3). On a 198-question causal benchmark the three reach 99.5%, 79%, and 75.8%, forming a deployment ladder above the strongest retrieval-augmented baseline at 48%. The DRR retriever has 110M parameters and trains per plant in ~17 seconds; after cross-plant transfer to a biologically distinct plant it still reaches 88%, while Method 2's static table cannot transfer. On a 60-question counterfactual benchmark only Method 3 handles queries about what happens after an intervention: +16.3 pp over Method 2, paired 95% CI [+7.1, +26.4] pp, with 100% on the timescale and operating-regime categories. On the AI2 Reasoning Challenge (ARC) with an OpenBookQA fact corpus, the same selective-retrieval mechanism reaches 79% versus unconstrained Llama-3.1-8B 76% and full-injection 74%, a +3 pp out-of-domain replication that argues against a result specific to wastewater treatment. We provide the first single-simulator comparison of live tool-use, static parameter injection, and learned numerical-parameter retrieval for industrial causal question answering.
Comments20 pages, 2 figures, 8 tables. Preprint submitted to Elsevier