发表机构
Beijing University of Posts and Telecommunications; Sinopec Engineering Incorporation(北京邮电大学; 中石化工程有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SkillRefine通过利用文档与实践差距、多源证据整合及执行反馈验证,为LLM代理在炼油规划软件中归纳技能,显著提升复杂多表协调任务性能。
AI 中文摘要
操作诸如AspenTech PIMS(过程工业建模系统)之类的工业规划软件,要求LLM代理结合结构知识、来自专家记录的程序性知识以及仅在执行过程中显现的约束。这些证据来源是异构的,且各自不完整:文档描述了表和接口,但省略了任务级协调;专家CASE记录暴露了多表修改模式,但没有明确的模式基础;执行反馈仅在计划执行时揭示潜在约束。我们提出了\ extsc{SkillRefine},一个利用文档化结构与专家实践之间的文档-实践差距来提取候选协调模式、根据表定义和COM规范对其进行基础化,并将其编译成具有来源的逐步披露的技能包框架。该库随后通过无标签合规性筛选、基于oracle的表、行、列和值维度的匹配分解,以及信号条件轨迹归因进行局部修复来精炼。我们在PIMS-Bench上评估了\ extsc{SkillRefine},这是一个基于两个AspenTech PIMS演示模型构建的基准,使用不相交的构建和保留测试任务。在保留测试集上,\ extsc{SkillRefine}在四个LLM骨干上实现了绝对组件匹配F1增益14%至30%,在复杂的多表协调任务上增益最大。
英文摘要
Operating industrial planning software such as AspenTech PIMS (Process Industry Modeling System) requires an LLM agent to combine structural knowledge, procedural knowledge from expert records, and constraints revealed only during execution. These evidence sources are heterogeneous and individually incomplete: documentation describes tables and interfaces but omits task-level coordination, expert CASE records expose multi-table modification patterns without explicit schema grounding, and execution feedback reveals latent constraints only when a plan is executed. We present \textsc{SkillRefine}, a framework that exploits the Documentation--Practice Gap between documented structure and expert practice to extract candidate coordination patterns, ground them against table definitions and COM specifications, and compile them into progressively disclosed skill packages with provenance. The library is then refined through label-free compliance screening, oracle-based match decomposition over table, row, column, and value dimensions, and signal-conditioned trajectory attribution for localized repair. We evaluate \textsc{SkillRefine} on PIMS-Bench, a benchmark built from two AspenTech PIMS demonstration models, using disjoint construction and held-out test tasks. On the held-out test set, \textsc{SkillRefine} achieves absolute component match F1 gains of 14\%--30\% across four LLM backbones, with the largest gains on complex multi-table coordination tasks.
Comments8 pages, 4 figures, 2 tables