AgentPProf:面向长时程AI智能体的语义剖析器
AgentPProf: Semantic Profiler for Long Horizon AI Agents
另 1 家 · 查看机构详情
- UC Santa Cruz(加州大学圣克鲁兹分校)
- Eunomia Labs(Eunomia 实验室)
- HKUST(香港科技大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
AgentPProf提出语义操作栈模型和递归操作分割,将智能体轨迹聚合成pprof兼容剖析文件,实现长时程任务的资源归因、问题定位和成本优化,在基准上显著提升性能。
中文摘要 AI 辅助
AI智能体越来越多地编排与用户、工具和系统资源的长时间活动,持续数天乃至数周。为提升智能体的质量、安全性和成本效率,开发者需要确定故障发生在何处、哪些因素触发不安全影响、哪些任务消耗最多预算,并优化这些任务。在系统软件中,剖析通过聚合资源消耗并将其归因于负责的代码路径来识别热点,从而回答类似问题。然而,现有的智能体可观测性工具侧重于单次执行的调试和追踪,而非跨运行、长期剖析,使得这些问题难以规模化回答。智能体可观测性需要剖析,而不仅仅是调试,但剖析智能体具有挑战性:责任实体是任务意图(如诊断认证、比较分支)而非代码路径,且缺乏用于聚合的稳定标识符。我们提出一种语义操作栈模型,将剖析适配到智能体轨迹。统一操作表示所有活动,操作栈替代运行时调用栈,实现不同粒度下的层次归因。我们观察到智能体的任务占据连续区间并可分解为子任务,因此引入递归操作分割,在任务边界递归分割轨迹。AgentPProf是一个剖析器,将智能体轨迹聚合成兼容pprof的剖析文件,支持火焰图可视化和分析。AgentPProf在CodeTraceBench上对人工标注达到0.764的B^3 F1分数。在三个问题定位基准上,剖析将MAP提升最多56%,表明其能有效归因资源、定位问题,并以实际剖析成本帮助优化令牌成本。AgentPProf可在该https URL获取。
英文摘要
AI agents increasingly orchestrate long-running activities with users, tools, and system resources for days and weeks. To improve agent quality, safety, and cost efficiency, developers need to determine where failures happen, what triggers unsafe effects, and which tasks consume the most budget, then optimize those tasks. In systems software, profiling answers similar questions by aggregating resource consumption and attributing it to responsible code paths to identify hotspots. Yet existing agent observability tools focus on per-execution debugging and tracing rather than cross-run, long term profiling, making these questions difficult to answer at scale. Agent observability needs profiling, not only debugging, but profiling agents is challenging: the responsible entities are task intent like diagnose authentication, compare branches rather than code paths, and lack stable identifiers for aggregation. We propose a semantic operation stack model that adapts profiling to agent trajectories. Uniform operations represent all activities, and operation stacks replace the runtime call stack, enabling hierarchical attribution at different granularities. We observe that an agent's task occupies a contiguous span and decomposes into subtasks, so we introduce recursive operation segmentation, which recursively splits trajectories at task boundaries. AgentPProf is a profiler that aggregates agent trajectories into pprof-compatible profiles, enabling flame graph visualization and analysis. AgentPProf reaches 0.764 $B^3$ F1 against human annotations on CodeTraceBench. On three problem-localization benchmarks, the profile raises MAP by up to 56%, demonstrating that it effectively attributes resources, locates problems, and helps optimize token cost at practical profiling cost. AgentPProf is available at https://github.com/eunomia-bpf/agentsight.