人类与人工智能协作的多线任务调整:基于本地大语言模型与数字孪生
From Human Intent to Strategy Decisions: A Review Framework Integrating Large Language Models and Digital Twins
浏览论文内容
中文总结 AI 辅助
本研究提出一种结合本地大语言模型、数字孪生和人类决策的多线任务调整系统,通过提议-验证-决定工作流实现可追溯的策略审查,实验显示分阶段筛选有效,但需进一步测试验证泛化性。
中文摘要 AI 辅助
自动化系统必须适应不断变化的任务、设备状态和人员配置条件,同时为人工审查提供证据。本研究提出了一种多线任务调整系统,该系统集成了本地大语言模型、数字孪生和人类决策。一个“提议-验证-决定”工作流将操作员意图转化为结构化需求,生成一组有界候选策略,并检查语义、仿真执行和操作约束。关联记录保持了从请求到验证证据和决策的可追溯性。使用四条虚拟手术器械分拣线评估了三十条固定测试记录:其中28条评估工作流,2条评估模型生成。十八条工作流案例符合预期;自主策略工作流成功率为3/10,正确拒绝无效输入为7/8。所有四个通过先前检查、产生完整证据并到达最终工程审查(CP6)的案例均通过了该审查。加上对未通过吞吐量约束的策略的正确阻止,这支持了在所测试环境中分阶段筛选和确认的有效性。八条仿真证据记录的平均放置验证通过率为97.50%。排除启动时间后,首次可审查响应和仿真验证的平均时间分别为12.94秒和164.39秒。剩余失败涉及语义扭曲、证据不完整和遗漏无效输入。结果表明了一个可追溯的策略审查工作流,但并未确立整体可靠性或长期稳定性。需要更广泛的测试和物理评估以评估泛化性。
英文摘要
Automation systems must adapt to changing tasks, equipment states, and staffing conditions while providing evidence for human review. This study presents an intent-review framework integrating large language models, digital twins, and human decision-making, implemented for multi-line task adjustment. A Propose-Verify-Decide workflow translates operator intent into structured requirements, generates a bounded set of candidate strategies, and checks semantics, simulation execution, and operational constraints. Linked records preserve traceability from requests to verification evidence and decisions. Thirty fixed test records were evaluated using four virtual surgical-instrument sorting lines: 28 assessed the workflow and two assessed model generation. Eighteen workflow cases met expectations (64.29%); three of ten strategy-workflow requests completed autonomously (30%); and seven of eight invalid inputs were correctly intercepted (87.5%). Analysis by anomaly source showed that the framework rejected two candidate batches that failed throughput constraints and paused recommendations in two cases with incomplete evidence. A mistranslated priority rule nevertheless reached simulation. These findings demonstrate handling of specific invalid inputs, throughput violations, and evidence gaps, and identify consistency between original intent and execution specifications as a key improvement area. The effect on overall erroneous decisions remains to be evaluated through matched comparisons, together with the influence of digital-twin fidelity using independent physical references.
发表机构
- Department of Mechanical Engineering, National Central University(国立中央大学机械工程系)
机构由 AI 辅助整理,请以论文原文为准。