arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07148cs.AI

面向智能制造中大型语言模型增强的多智能体强化学习(MARL)中心参考架构

A MARL Centered Reference Architecture for Large Language Model Augmentation in Smart Manufacturing

Fouad Bahrpeyma, Dirk Reichelt

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对智能制造自适应控制的耦合需求,提出以MARL为中心的三层参考架构,探讨LLM在MARL中的附着点,明确传统MARL与LLM组件的适配场景,为相关研究提供结构化框架。

中文摘要 AI 辅助

现代智能制造对自适应控制提出了六项耦合需求:具有全局影响的局部决策、部分可观测性、非平稳性、具备长时程效应的反射速度响应、延迟且分散的结果,以及难以显式建模的动态特性。合作多智能体强化学习(MARL)以集中式训练与分布式执行的分布式部分可观测马尔可夫决策过程(Dec-POMDP)形式提出,是适配这些需求的特别自然的范式。本文采用以MARL为中心的范围,探讨大型语言模型(LLM)应在何处增强、与该协调核心交互、训练,或在最强竞争场景中替代该协调核心。本文通过四个LLM附着点组织文献分类:策略、奖励设计、智能体间通信以及分层规划。条件能力概况区分了原生机制、报告性能、形式保证和工程成熟度,部署就绪度分析确定了每个角色背后的证据。这些阶段产生了主要贡献:一个基于证据的三层以MARL为中心的参考架构,用于语义推理、自适应合作控制和独立保证执行。LLM增强型Dec-POMDP是该架构的描述性比较符号,记录了四个附着选择,而不引入新的决策过程类或算法。根据所审查的证据,传统MARL更适合在特定任务训练后进行频繁、结构化的分布式协调,而LLM组件在语义解释、奖励起草、人机交互和较慢的监督规划方面具有前景。当前仅基于LLM的制造控制器尚未在严格的实时、分布式、安全关键控制方面建立等效性;该结论受可用证据的限制,并未断言其不可能实现。

英文摘要

Modern manufacturing imposes six coupled demands on adaptive control: local decisions with global consequences, partial observability, nonstationarity, reflex speed response with long horizon effects, delayed and diffuse outcomes, and dynamics that resist explicit modeling. Cooperative multiagent reinforcement learning (MARL), posed as a Dec-POMDP under centralized training with decentralized execution, is a particularly natural formalism for these demands. This paper adopts a MARL centered scope and asks where large language models (LLMs) should augment, interface with, train, or, in the strongest competitive case, replace that coordination core. A taxonomy organizes the literature through four LLM attachment points: policy, reward design, communication between agents, and hierarchical planning. A conditional capability profile separates native mechanism, reported performance, formal guarantee, and engineering maturity, and a deployment readiness analysis identifies the evidence behind each role. These stages yield the principal contribution: a three layer MARL centered reference architecture, grounded in evidence, for semantic reasoning, adaptive cooperative control, and independently assured execution. The LLM-Augmented Dec-POMDP is a descriptive comparative notation for that architecture, recording four attachment choices without introducing a new decision process class or algorithm. Under the reviewed evidence, conventional MARL is better suited to frequent, structured, decentralized coordination after task specific training, whereas LLM components are promising for semantic interpretation, reward drafting, human interaction, and slower supervisory planning. Current LLM only manufacturing controllers do not yet establish equivalence for strict real time, decentralized, safety critical control; this conclusion is bounded by the available evidence and does not assert impossibility.

↑