arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04265cs.MAcs.AIcs.SYeess.SY

网络物理系统中大语言模型智能体规划策略的战略评估

Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems

J. de Curtò, I. de Zarzà

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对网络物理系统,构建含40个异构产消者的智能电网基准,评估4种LLM规划架构,发现架构影响结果,应用截止可行性可显著降低遗憾。

中文摘要 AI 辅助

对大语言模型(LLM)规划智能体的评估大多仅关注任务是否成功或是否遵循已声明的计划。在战略网络物理系统中,更重要的问题是规划架构在自主参与者做出响应且物理过程约束结果后是否仍适用。我们引入了一个受物理约束的受控基准,该基准围绕规划诱导的控制轨迹构建,即执行架构作用于其他智能体和物理过程的有序规划操作与指令。该基准在拥有40个异构产消者(prosumer)及独立模拟径向馈线的智能电网需求响应系统中实现了预定义、顺序式、分层式和搜索式执行器。LLM被限制为类型化策略声明和短操作消息,而调度构建、产消者动态及潮流则保留为显式代码。该协议采用配对强制模式反事实、通用随机响应抽样和事件级截止可行性。研究得出三个关键性质:架构会实质性改变结果,强制搜索在全部5个基准种子中均为最优解;执行保真度不止需要模式一致,目标替换在达成100%一致的同时,电压缺口却增加了2.68倍;在包含144个场景、576个回合的基准库中,4种架构中有3种存在可行最优解;预指定的应力保留岭回归的平均遗憾为90.7(95%置信区间[73.8,108.6]),且未表现出优于固定顺序的性能;在质量预测前应用已知截止可行性可将遗憾降至29.0,较固定顺序提升61.1;全可行消融实验未优于固定搜索,表明剩余挑战在于可行范围内的质量选择;五模型扩展将应力条件、状态盲和不变声明器区分开,延迟尾部表明实时可行性应被视为概率性处理。

英文摘要

LLM-agent evaluations commonly measure task success or agreement with a declared plan. In strategic cyber-physical systems, an architecture must also remain appropriate after autonomous participants respond and physics constrains outcomes. We introduce a controlled benchmark of planning-induced control trajectories: ordered planning operations and directives linking execution architecture to strategic response and physical consequences. Four coded executors (predefined, sequential, hierarchical, and search) control demand response for 40 prosumers on a radial feeder. The LLM declares or advises typed policies and mediates communication; schedules, base prosumer dynamics, stochastic actions, and power flow remain explicit code. Paired forced-mode counterfactuals, exact-prompt caching, common response draws with separate randomness streams, critic isolation, and event-level feasibility isolate comparisons. The Llama-3.3-70B experiments on this feeder distinguish three properties. First, forced search is the oracle in all five baseline seeds under the specified objective. Second, injected objective substitution preserves mode agreement at 1.0 while increasing cumulative voltage shortfall by 2.68x. Third, the 144-scenario, 576-episode factorial bank, using three repeated seeds, contains feasible oracles from predefined, sequential, and search. The prespecified stress-held-out ridge has mean regret 90.7 and no observed value over fixed sequential. A post-hoc constraint-aware analysis reduces regret to 29.0; a simple deadline rule attains 28.7, so this gain does not establish a learning advantage. An all-feasible ablation does not improve over fixed search. These are simulation-internal, descriptive comparisons. A five-model, 300-declaration extension tests interface behaviour, not cross-backbone physical rankings; shared-endpoint latency tails motivate probabilistic live feasibility.

发表机构

  • BARCELONA Supercomputing Center(巴塞罗那超级计算中心)
  • LUXEMBOURG Institute of Science and Technology(卢森堡科学技术学院)

机构由 AI 辅助整理,请以论文原文为准。

↑