用于工业过程操作决策的大语言模型引导型上下文动作评估(LCAE)
LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes
- College of Information Science and Engineering, Northeastern University(东北大学信息科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出LCAE方法,用LLM标准化工业文档为关系基础,结合历史调节关系强度,让文档语义融入策略学习,以优化工业过程操作决策的动作评估。
AI中文摘要:
工业演员-评论家方法通常将连续动作表示为匿名数值坐标,因此必须从有限的交互中学习每个动作影响哪些过程变量、影响方向以及延迟时间。固定工业文档已描述了部分此类关系,但其开放式文本陈述既未反映当前运行工况,也无法直接适配数值策略。本文提出用于工业过程操作决策的大语言模型引导型上下文动作评估(LCAE),该方法在训练前使用大语言模型将固定文档标准化为冻结的动作-观测-方向-延迟关系基础;近期的数值动作-响应历史会调节每个关系的当前强度,而被评估动作会在同一基础上形成状态条件非线性动作效应场。评论家通过该场评估动作,演员利用相同关系增益生成动作,使文档语义成为最大熵策略学习的一部分。在训练或部署期间,LLM(大语言模型)和嵌入模型均不在线运行;部署的策略仅使用冻结的语义工件和可见的数值历史。该方法提出可证伪假设:当文档记录的关系正确且近期历史反映其上下文强度时,此动作表示应比原始动作坐标提供更有用的决策偏差。
英文摘要:
Industrial actor--critic methods usually represent continuous actions as anonymous numerical coordinates. They must therefore learn from limited interactions which process variables each action affects, in which direction, and after what delay. Fixed industrial documents already describe part of these relations, but their open-text statements neither represent the current operating condition nor directly fit a numerical policy. This article presents LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes (LCAE), which uses a large language model before training to normalize fixed documents into a frozen action--observation--direction--delay relation basis. Recent numerical action--response history then modulates the current strength of each relation, while the evaluated action forms a state-conditioned nonlinear action-effect field in the same basis. The critic evaluates actions through this field, and the actor uses the same relation gains to generate actions, making document semantics part of maximum-entropy policy learning. Neither the LLM nor the embedding model runs online during training or deployment; the deployed policy uses only frozen semantic artifacts and visible numerical history. The method states a falsifiable hypothesis: when documented relations are correct and recent history reflects their contextual strength, this action representation should provide a more useful decision bias than raw action coordinates.