发表机构
National University of Singapore; Fudan University; The University of Sydney(新加坡国立大学; 复旦大学; 悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对演化环境中仅凭结果评估智能体的局限,提出过程感知基准LiveMACEBench,利用实时金融市场诊断LLM智能体能力,揭示结果与能力间的显著差距。
AI 中文摘要
仅依据结果来评估智能体可能会掩盖产生这些结果的能力。这一问题在演化环境中尤为突出,因为结果反映的是智能体行为与不断变化的外部条件之间的闭环交互。我们引入了LiveMACEBench,一个过程感知基准,利用实时金融市场作为自然演化的测试平台,用于评估持久化LLM智能体。五个前沿LLM在匹配的工具使用、持久记忆、规则遵循和多智能体协作配置下沿连续轨迹运行。我们通过实际结果以及从完整决策轨迹中推导出的机制特定诊断来评估它们。在30天的实时评估中,我们发现了一个显著的结果-能力差距:实际回报往往与能力特定测量存在分歧,且相似的结果可能源于截然不同的机制使用模式。轨迹级诊断进一步揭示了不同能力间的明显瓶颈,表明机制访问、有效机制使用和下游性能并非可互换的智能体能力度量。LiveMACEBench使这一区分变得可测量,将实时市场从性能排行榜转变为智能体能力的诊断环境。
英文摘要
Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability