测量语言模型智能体的跨任务行为一致性
Measuring Cross-Task Behavioral Consistency in Language Model Agents
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
- Mantis AI Research(螳螂人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出行为一致性指标(BCM),通过约9000条软件工程任务轨迹发现语言模型智能体的跨任务与任务内一致性是可分化的不同维度,且一致性与成功率无关,BCM可补充结果指标评估智能体。
AI中文摘要:
智能体评估几乎完全依赖成功率等结果指标,这些指标仅能判断智能体是否成功,却无法衡量其行为的一致性。本文提出跨任务行为一致性是一种独特且可测量的属性,并引入行为一致性指标(Behavioral Consistency Metric,BCM)对其进行量化。BCM 训练一个模型,从智能体执行轨迹的行为特征中预测任务成功情况,推导每个轨迹的特征归因向量,并测量智能体系统内这些向量的平均成对相似度。在来自六个语言模型智能体、涉及软件工程任务的约 9000 条轨迹中,核心发现是跨任务一致性和任务内一致性是可分化的不同维度:部分系统在局部可复现,即同一任务的重复尝试中行为相似,但在全局上分散,不同任务间无稳定策略;而其他系统在两个尺度上均保持一致。过往研究仅测量同一任务的可复现性,无法观察到这种分化。进一步研究发现,一致性无法简化为成功率,因为成功率相当的系统在一致性上可能存在显著差异;且在保持任务难度一致的任务内控制条件下,闭源与开源模型的一致性差距依然存在。本文将 BCM 定位为补充结果指标的过程级可靠性信号,并明确了其适用的有效条件。
英文摘要:
Agent evaluation relies almost entirely on outcome metrics such as success rate, which capture whether an agent succeeds but not how consistently it behaves. We argue that behavioral consistency across tasks is a distinct and measurable property, and we introduce the Behavioral Consistency Metric (BCM) to quantify it. BCM trains a model to predict task success from behavioral features of agent execution traces, derives a per-trajectory feature-attribution vector, and measures the mean pairwise similarity of these vectors within an agent system. Across roughly 9,000 trajectories from six language model agents on software engineering tasks, our central finding is that cross-task and within-task consistency are distinct axes that can diverge: some systems are locally reproducible, behaving similarly on repeated attempts at one task, yet globally fragmented, with no stable strategy across different tasks, while others are consistent at both scales. Prior work measures only same-task reproducibility and so cannot observe this separation. We further find that consistency is not reducible to success rate, since systems with comparable success can differ sharply in consistency, and that the frontier-versus-open-source consistency gap persists under a within-task control that holds task difficulty constant. We position BCM as a process-level reliability signal that complements outcome metrics, and we are explicit about the conditions under which it is meaningful.