发表机构
College of Science and Engineering, Hamad Bin Khalifa University(科学与工程学院,哈马德·本·哈利法大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出将大型语言模型作为点对点能源双向拍卖中的策略性投标智能体,相比ε-贪心多臂老虎机和随机投标,LLM策略在交易量和跨环境泛化上更优,但同质LLM群体可能导致市场剩余分配失衡。
AI 中文摘要
点对点(P2P)能源交易市场依赖双向拍卖机制来匹配智能电网配电网络中的产消者和消费者。然而,此类市场中投标智能体的策略性行为尚未得到充分探索,尤其是在具有有限理性的重复博弈环境中。本文提出了一种新颖框架,将大型语言模型(LLM)作为基于推理的投标智能体集成到重复的P2P能源双向拍卖中。我们比较了三种投标策略的性能:随机投标、一种ε-贪心多臂老虎机(MAB)方法和一种基于LLM的策略。仿真结果表明,与ε-贪心MAB和随机投标基线相比,基于LLM的策略在最初几轮中实现了更高的已清算交易量,消除了统计算法在收敛到有效价格臂之前固有的探索预热期。重要的是,当在与学习期间不同的环境中进行测试时,ε-贪心策略的性能显著下降,而基于LLM的投标策略继续实现更高的剩余和成功交易。然而,LLM的高级上下文推理也引发了一个重要的市场动态。在同质的LLM智能体群体中,卖方越来越多地利用买方的理性外部选择权,将清算价格推高至纳什均衡之上,导致剩余分配逐渐向卖方倾斜,且在观察时间范围内未收敛。
英文摘要
Peer-to-peer (P2P) energy trading markets rely on double auction mechanisms to match prosumers and consumers in smart grid distribution networks. However, the strategic behavior of bidding agents in such markets remains not fully explored, particularly in repeated settings with bounded rationality. This paper proposes a novel framework that integrates large language models (LLMs) as reasoning-driven bidding agents in repeated P2P energy double auctions. We compare the performance of three bidding strategies: random bidding, an $\varepsilon$-greedy multi-armed bandit (MAB) approach, and an LLM-based strategy. Simulation results show that the LLM-based strategy achieves superior cleared trading volume over the first episodes compared to $\varepsilon$-greedy MAB and random bidding baselines, eliminating the exploration burn-in period that statistical learning algorithms inherently require before converging to productive price arms. Importantly, when tested in an environment different from the one used during learning, the performance of the $\varepsilon$-greedy strategy drops significantly, while that of the LLM-based bidding strategy continue to achieve higher surplus and successful trades. Nevertheless, the LLM's advanced contextual reasoning also gives rise to an important market dynamic. In a homogeneous population of LLM agents, sellers increasingly exploit buyers' rational outside options to drive clearing prices above the Nash equilibrium, resulting in a progressively more asymmetric allocation of surplus in favor of sellers that does not converge within the observed time horizon.
CommentsAccepted at IEEE IECON 2026