arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SolarChain-Eval:去中心化能源市场中可信经济主体的物理约束基准测试

SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets

Shilin Ou, Yifan Xu, Luyao Zhang

arXiv 2607.08681首次发表:更新:

发表机构

Duke Kunshan University(杜克昆山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对去中心化能源市场中智能代理评估需兼顾任务性能与可信度的问题,提出物理约束基准测试SolarChain-Eval,将市场治理建模为马尔可夫决策过程,从多维度评估策略,纳入大语言模型规划器/审计器层,实验揭示效用与安全权衡及相关问题。

AI 中文摘要

随着智能代理人工智能系统越来越多地应用于网络物理环境,其评估需要同时考量任务性能和可信度。在去中心化能源市场中,自主代理虽能提升市场效用,但也可能利用无效物理数据、制造人为流动性并做出不稳定的治理决策。因此,我们提出了SolarChain-Eval,这是一个用于评估可信经济主体的物理约束基准测试。它将市场治理表述为一个与Gymnasium兼容的马尔可夫决策过程,代理每小时做出决策。SolarChain-Eval从多个维度评估每个策略,包括市场效用、物理安全性、滑点、行动平滑性、空间公平性和可审计性。为支持智能代理评估,SolarChain-Eval纳入了基于大语言模型的规划器/审计器层。规划器定义情节级行动边界和审计规则,审计器审查并修订高风险行动。所有干预都通过结构化日志记录,包括触发信号、提议行动、修订行动和审计理由。对静态、随机、近视、强化学习和强化学习+大语言模型策略的实验揭示了效用与安全之间明显的权衡。强化学习代理提高了市场效用,但仍可能产生不安全行为。去除物理惩罚后,追求奖励最大化的代理会利用无效发电并增加人为流动性。大语言模型规划器/审计器提高了可审计性并减轻了特定风险,但无法完全弥补错误指定的奖励函数。这些结果表明,可信的智能代理人工智能评估既需要物理约束,也需要透明的干预痕迹。我们在GitHub上以开放访问的方式发布数据和代码以确保可重复性。

英文摘要

As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performance and trustworthiness. In decentralized energy markets, autonomous agents may improve market utility, but may also exploit invalid physical data, create artificial liquidity, and produce unstable governance decisions. Therefore, we propose SolarChain-Eval, a physics-constrained benchmark for evaluating trustworthy economic agents. It formulates market governance as a Gymnasium-compatible Markov Decision Process, where agents make hourly decisions. SolarChain-Eval evaluates each policy across multiple dimensions, including market utility, physical safety, slippage, action smoothness, spatial fairness, and auditability. To support agentic evaluation, SolarChain-Eval incorporates an LLM-based Planner/Auditor layer. The Planner defines episode-level action bounds and audit rules, while the Auditor reviews and revises high-risk actions. All interventions are recorded through structured logs, including trigger signals, proposed actions, revised actions, and audit rationales. Experiments with static, random, myopic, RL, and RL+LLM policies reveal a clear utility-safety trade-off. RL agents improve market utility but can still produce unsafe behavior. When the physics penalty is removed, reward-maximizing agents exploit invalid generation and increase artificial liquidity. The LLM Planner/Auditor improves auditability and mitigates selected risks, but it cannot fully compensate for a misspecified reward function. These results indicate that trustworthy agentic AI evaluation requires both physical constraints and transparent intervention traces. We release data and code as open access on GitHub for replicability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑