大语言模型智能体多轮一致性的评估:生存分析与失败理由分类学
Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy
- Carleton University(卡尔顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过20步延迟满足实验和84,540条轨迹,结合生存分析与理由分类学,评估了LLM智能体在多轮交互中的时间一致性,发现深思与矛盾率相关,并揭示了模型特定的失败指纹。
AI中文摘要:
大语言模型(LLM)智能体在孤立任务上可能表现良好,但在长时间交互中会逐渐出现不一致。我们在一个受延迟满足研究启发的受控20步多智能体环境中评估时间一致性。在每一步中,智能体选择继续延迟奖励或立即领取奖励(终止回合)。通过对社会可见性(私有vs公共)、人格压力源和深思策略的全因子操纵,我们运行了涵盖8个模型家族的84,540条轨迹。将首次领取奖励视为时间至事件结果,我们估计了Kaplan-Meier生存曲线并拟合离散时间风险回归,以量化实验因素如何随时间改变失败风险。然后,为了分析与失败相关的理由和语言模式,我们利用LLM辅助标注并配合人工审计(κ=0.83),从选择终止回合的智能体的13,780条深思轨迹中构建了一个七类别分类体系。理由概况随时间和情境系统性变化:早期失败更多由冲动驱动,后期失败更多以疲劳和成本效益为框架,而公共环境增加了规范导向的辩解。我们还发现深思与不一致之间的关联:在失败中,较长的深思与更高的理由内部矛盾率(同时出现支持延迟和支持领取的陈述)相关,这挑战了更多推理文本意味着更高一致性的假设。综合来看,生存分析和理由分析揭示了不同的时间可靠性机制和模型特定的“失败指纹”,为诊断多轮智能体行为中的不一致性提供了评估视角。
英文摘要:
Large language model (LLM) agents may perform well on isolated tasks yet drift into inconsistency over extended interaction. We evaluate temporal consistency in a controlled 20-step multi-agent setting inspired by delayed-gratification studies. At each step, an agent chooses between continuing to delay a reward or claiming it immediately (terminating the episode). Across a full-factorial manipulation of social visibility (private vs public), persona stressors, and deliberation policy, we run 84,540 trajectories spanning 8 model families. Treating the first reward-claim as a time-to-event outcome, we estimate Kaplan-Meier survival curves and fit discrete-time hazard regression to quantify how experimental factors shift failure risk over time. Then, to analyze rationales and language patterns associated with failure, we build a seven-category taxonomy from 13,780 deliberation traces from agents who choose to terminate the episode, using an LLM-assisted labeling paired with human audit ($κ=0.83$). Rationale profiles change systematically with time and context: early failures are more impulse-driven, later failures more fatigue- and cost-benefit-framed, while public settings increase norm-oriented justifications. We also find a deliberation-inconsistency association: among failures, longer deliberation correlates with higher rates of intra-rationale contradiction (simultaneous pro-delay and pro-claim statements), challenging the assumption that more reasoning text implies greater consistency. Together, the survival and rationale analyses reveal distinct temporal reliability regimes and model-specific "failure fingerprints", offering an evaluation lens for diagnosing inconsistency in multi-turn agent behavior.