大语言模型智能体轨迹中的成本-效用对齐:分析、归因、诊断、适配与评估
Cost-Utility Alignment in LLM Agent Trajectories:Profiling,Attribution,Diagnosis,Adaptation,and Evaluation
浏览论文内容
中文总结 AI 辅助
本研究提出以轨迹为中心的成本-效用对齐框架,通过五阶段分析解决LLM智能体资源消耗与任务贡献不匹配问题,为资源感知的智能体设计部署提供结构化基础。
中文摘要 AI 辅助
大语言模型(LLM)智能体通过多步骤轨迹执行任务,这些轨迹会累积token、延迟、货币费用及环境风险等成本,仅在整体任务层面产生效用。现有研究仅孤立地处理推理优化、智能体能力或评估问题,导致从业者缺乏原则性工具来判断轨迹的资源消耗是否与其任务贡献相匹配。为填补这一空白,我们开发了一种以轨迹为中心的成本-效用对齐框架,将资源消耗与任务贡献视为同一执行过程中的双重账本,围绕五个分析阶段组织:成本分析、效用归因、错配诊断、针对性适配及评估。效用归因是该框架的核心:它不依赖整体结果,而是按证据强度对贡献方法进行组织,从过程代理、信息依赖到反事实重放,提供支撑诊断并指导适配的因果证据。利用该框架,我们分析了近期的智能体系统、归因方法及覆盖效率、可靠性和经济价值的评估协议,以及五种错配形式,涵盖认知与上下文使用、外部交互、恢复循环控制、资源-能力分配及多智能体协调,还有它们的针对性适配。最终形成了连接智能体执行成本侧与效用侧的闭环分析,为感知资源的智能体设计与部署提供结构化基础。
英文摘要
LLM agents execute tasks through multi-step trajectories that accumulate cost in tokens, latency, monetary fees, and environmental risk while producing utility only at the aggregate task level. Prior surveys address inference optimization, agent capabilities, or evaluation in isolation, leaving practitioners without principled tools to determine whether a trajectory's resource expenditure is justified by its task contribution. We address this gap by developing a trajectory-centric cost-utility alignment framework that treats resource consumption and task contribution as dual ledgers over the same execution, organized around five analytical stages: cost profiling, utility attribution, misalignment diagnosis, targeted adaptation, and evaluation. Utility attribution is central to this structure: rather than relying on aggregate outcomes, it organizes contribution methods by evidential strength, from process proxies and information dependency to counterfactual replay, supplying the causal evidence that grounds diagnosis and guides adaptation. Using this framework, we analyze recent agent systems, attribution methods, and evaluation protocols covering efficiency, reliability, and economic value, as well as five forms of misalignment spanning cognitive and context use, external interaction, recovery-loop control, resource-capability allocation, and multi-agent coordination, together with their targeted adaptations. The result is a closed analytical loop connecting the cost side of agent execution to its utility side, providing a structured basis for resource-aware agent design and deployment.