发表机构
Carleton University(卡尔顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一个多智能体微基准,模拟棉花糖实验,通过因子化操纵社会背景、人格和工具使用政策,用生存分析揭示LLM智能体在长视界中的延迟满足行为,发现隔离降低风险而强制提问增加风险,消融实验提升完成率。
AI 中文摘要
大语言模型(LLM)越来越多地被部署为多轮智能体,必须在长时间交互中维持目标、使用工具并适应其他智能体。然而,现有研究缺乏可审计的、多轮、多因素的实验来量化LLM在明确约束下的行为,且缺乏揭示行为在长视界中如何展开的时间分辨统计。为填补这一空白,我们开发了一个受斯坦福棉花糖实验启发的多智能体微基准:ReAct智能体在每分钟的时间步上运行,使用“提出问题”工具,并受每步预算限制,同时我们因子化地操纵社会背景(广播与隔离)、人格(年龄、享乐驱动)和元认知策略(强制与可选工具使用)。我们使用Kaplan-Meier(KM)生存曲线和离散时间风险模型,在长风险视界内分析64个单元中19,200条智能体轨迹的结果。行为显示出明显的早期“进食”冲动,仅75.9%的智能体坚持到最后。在离散时间风险模型中,隔离相对于广播降低了每分钟风险,而必须使用自我提问的政策增加了风险。平均而言,智能体提出约7.12个问题,并在约6%的分钟中达到每步预算。在广播条件下,提问下降速度快于隔离条件。消融实验表明,去除享乐驱动和/或人格年龄可提高生存率和完成率,缩小广播/隔离差距,但保持必须与可选顺序不变。组合消融(无享乐+无人格年龄)产生最高完成率(接近1.0)。这些结果确立了延迟满足作为一个紧凑的多轮交互基准,能够捕捉LLM智能体中的社会传染和工具使用动态,为分析长视界、多智能体行为提供了可复现的测试平台和统计数据。
英文摘要
Large language models (LLMs) are increasingly deployed as multi-turn agents that must sustain goals, use tools, and adapt to other agents over extended interactions. However, existing research lacks auditable, multi-turn, multi-factorial experiments that quantify LLM behavior under explicit constraints, with time-resolved statistics that reveal how behavior unfolds over long horizons. To address this gap, we develop a multi-agent micro-benchmark inspired by the Stanford marshmallow experiment: ReAct agents operate minute-by-minute with a "raise a question" tool under a per-step budget, while we factorially manipulate social context (broadcast vs. isolated), personas (age, hedonic drive), and metacognitive policy (mandatory vs. optional tool use). We analyze outcomes with Kaplan-Meier (KM) survival curves and discrete-time hazard models over a long risk horizon across 19,200 agent trajectories in 64 cells. Behavior shows a sharp early "eat" impulse, and only 75.9% of agents persist to the end. In a discrete-time hazard model, isolation reduces per-minute risk relative to broadcast, whereas a must-use self-questioning policy increases risk. On average, agents ask $\approx 7.12$ questions and hit the per-step budget in $\approx 6\%$ of minutes. Questioning declines faster under broadcast than isolation. Ablation experiments demonstrated that removing hedonic drive and/or persona age increases survival and completion, narrows the broadcast/isolated gap, but leaves the must vs. may ordering intact. The combined ablation (no hedonic + no persona age) yields the highest completion (approaching $1.0$). These results establish delay-of-gratification as a compact, multi-turn interaction benchmark that captures social contagion and tool-use dynamics in LLM agents, providing a reproducible testbed and statistics for analyzing long-horizon, multi-agent behavior.
CommentsAccepted as a poster at the NeurIPS 2025 Workshop on Multi-Turn Interactions in Large Language Models