AI 中文总结
研究大语言模型智能体在共享可再生能源储备时的协调失败问题,通过让四个同系列智能体自博弈,改变储备再生率,发现它们在需求超阈值时过度占用致自我挫败,实际消耗类似更不耐烦的开放访问基准,孤立响应评估会忽略此系统层面协调失败。
AI 中文摘要
大语言模型越来越多地被部署为能够规划、使用工具并长期行动的智能体。当它们共享持久资源,如计算池或能源储备时,一个智能体的决策会影响后续智能体面临的条件。我们在可再生能源共享池中研究这种协调失败。四个同系列的GPT、Gemini或Grok智能体作为电力生产者进行同质自博弈,被要求最大化运营连续性。在保持总剩余需求和决策协议不变的情况下,我们将共享能源储备的再生率从充足变为稀缺。当需求不超过可再生能源峰值替代量时,所有三个系列都能保护储备,但超过该阈值就会过度占用(所有九个精确的稀缺对比在霍尔姆校正后仍然显著;最大调整p值 = 4.87e - 5)。这种模式是自我挫败的:相同的群体在保护当前服务的同时破坏了未来服务。在更高的稀缺度(rho = 1.2)下,每个系列早期的总请求压力都超过了可再生能源峰值替代量,平均为该水平的1.21倍。平均轨迹在第5 - 7轮低于最大补给的储备水平。两个离线基准将最大化全组运营服务价值的社会规划者与开放访问进行比较,在开放访问中每个生产者最大化自身价值。在折扣因子gamma = 0.95时,两个基准在相同动态下都能维持储备。实际的消耗反而类似于更不耐烦的开放访问基准下的结果。因此,在公共轨迹层面,群体的行为就像不耐烦的优化者。这种系统层面的协调失败会被孤立响应评估所忽略。
英文摘要
LLMs are increasingly deployed as agents that plan, use tools, and act over time. When they share persistent resources, such as compute pools or energy reserves, decisions by one agent affect the conditions faced by later agents. We study this coordination failure in a renewable energy commons. Four same-family GPT, Gemini, or Grok agents act in homogeneous self-play as electricity prosumers, instructed to maximize operational continuity. Holding aggregate residual demand and the decision protocol fixed, we vary the regeneration rate of a shared energy reserve from abundance to scarcity. All three families preserve the reserve when demand does not exceed peak renewable replacement, but over-appropriate it beyond that threshold (all nine exact scarcity contrasts survive Holm correction; largest adjusted p = 4.87e-5). The pattern is self-defeating: the same populations protect current service while undermining future service. At higher scarcity (rho = 1.2), early aggregate request pressure exceeds peak renewable replacement in every family and averages 1.21 times that level. Mean trajectories fall below the reserve level of maximum replenishment by rounds 5-7. Two offline benchmarks compare a social planner maximizing group-wide operational-service value with open access, where each prosumer maximizes its own value. At a discount factor of gamma = 0.95, both benchmarks sustain the reserve under the same dynamics. Realized depletion instead resembles outcomes under a more impatient open-access benchmark. The populations therefore behave like impatient optimizers at the level of the public trajectory. This system-level alignment failure would be missed by isolated-response evaluation.