arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01466cs.AI

解析流:面向长视野智能体及其观察者的实时轨迹模型

Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers

Egor Pakhomov, Erik Nijkamp

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出实时轨迹模型,可将长视野智能体的轨迹折叠为类型化运行状态,经评估能降低观察者侧的令牌用量与成本、提升准确率,在智能体侧序列依赖任务中表现优于全上下文提示,相关成果与代码已公开。

中文摘要 AI 辅助

长视野智能体的轨迹超出了其两类使用者的承载能力:一类是监控运行过程的人类观察者,另一类是需将轨迹折叠回有限上下文的智能体自身。我们提出了一种实时轨迹模型,这是一种仅追加的事件账本,会被逐步折叠为类型化的运行状态,并编译为面向不同使用者的视图,随后针对确定性基准真值对两类使用者分别开展评估。针对观察者侧,以大语言模型(LLM)阅读器作为代理进行评估,编译后的视图回答监控问题时,阅读器所需输入令牌量比预算受限的单次调用读取原始轨迹少约14倍至15倍,成本低5倍至7倍,且准确率更高(0.85-0.87,对比原始轨迹的0.48)。由于这些问题是与视图模式协同设计的,我们将在模式覆盖范围内实现的令牌与成本降低视为可迁移的成果。针对智能体侧,在120个链接的序列依赖任务中,在每步状态中维护任务运行统计量的机制取得了全上下文提示失败的效果(在干净协议下为30/30,对比8/30,n=30,因与基准系统协同开发而标记为描述性结果);提示级草稿板在成本更低的情况下达到了与折叠操作相当的准确率,而双臂分解法将折叠操作的准确率归因于其确定性聚合,成本优势则归因于其紧凑性。折叠操作相较于更廉价替代方案的剩余价值在于可确定性审计,以及从同一状态为观察者提供服务。我们从观测到的失败中推导出11项轨迹折叠的候选需求,并以一个对顺序敏感的任务族界定了折叠不再起作用的边界。代码、基准、可复现的合成语料库以及所有工作台轨迹均已发布。

英文摘要

A long-horizon agent's trace outgrows both of its consumers: the human observer monitoring the run, and the agent itself, whose bounded context the trace must be folded back into. We present a live trace model, an append-only event ledger folded incrementally into typed run state and compiled into per-consumer views, and evaluate it for both consumers against deterministic ground truth. For the observer side, evaluated with an LLM reader as proxy, the compiled view answers monitoring questions using approximately 14x and 15x fewer input tokens (by reader) and at 5-7x lower cost than a budget-capped single-call reading of the raw trace, with higher accuracy (0.85-0.87 versus 0.48). Because the questions were co-designed with the view schema, we treat the token and cost reduction, conditional on schema coverage, as the transferable result. For the agent, on 120-link sequential-dependency tasks, mechanisms that maintain the task's running statistic in per-step state succeed where full-context prompting fails (30/30 versus 8/30 under a clean protocol, n=30, labeled descriptive owing to benchmark-system co-development); a prompt-level scratchpad matches the fold's accuracy at lower cost, and a two-arm decomposition attributes the fold's accuracy to its deterministic aggregate and its cost advantage to its compactness. The fold's remaining value over cheaper alternatives is deterministic auditability and serving the observer from the same state. We derive eleven candidate requirements for trace folding from observed failures and delimit them with an order-sensitive task family on which the fold ceases to help. Code, benchmarks, a regenerable synthetic corpus, and all workbench traces are released.

发表机构

  • Salesforce AI Research(Salesforce AI研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑