arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AstronOS:面向长周期智能体系统的统一执行模型与运行时

AstronOS: A Unified Execution Model and Runtime for Long-Horizon Agentic Systems

Zhenhang Nie, Gui Zheng, Xudong Sun, Tailong Zhu, Bin Zhang

arXiv 2608.16381首次发表:更新:

AI 中文总结

AstronOS是面向长周期智能体系统的统一执行模型与运行时,其介导的交接策略在三阶段任务中通过率远高于其他策略,虽时间成本更高但端到端表现更优。

AI 中文摘要

智能体系统通常围绕单一对话、模型调用或智能体实例来组织执行与状态,即便实际工作涉及大量调用和阶段。我们提出一种统一执行模型,该模型能在多次调用中维护工作项的持久身份与带版本的权威状态。每一步骤接收限定于特定状态版本和新素材的输入,结果仅在验证和记录后才推进状态。我们在AstronOS中实现了该模型的选定路径,涵盖Cases、Tasks和Scenario Packs,涉及中央与本地执行。我们对比了五种将既定软件版本更新计划带入全新模型会话的完整策略:重读原始素材、重放完整历史、确定性文本摘要、确定性JSON以及AstronOS运行时介导的交接。在所有五种策略下运行10个受控任务,每个任务重复3次,共150次执行。在单阶段参考系列中,各策略表现相似;在主要的三阶段A-C批次中,AstronOS在15次执行中有14次通过冻结评分器,而重读策略为0/15,完整历史重放为2/15,后续非交错摘要与JSON批次各为0/15。AstronOS在每次通过执行中按尝试次数计算的模型令牌成本更低,但每次尝试需要更多执行窗口时间。这些结果表明,在该基准中,完整的AstronOS条件与全新会话下更高的端到端通过率相关联,同时存在可测量的时间成本。

英文摘要

Agentic systems often organize execution and state around a single conversation, model invocation, or agent instance, even when real work spans many calls and stages. We introduce a unified execution model that maintains a work item's persistent identity and versioned authoritative state across calls. Each step receives input scoped to a specific state version and new material; a result advances state only after validation and recording. We implement selected paths of this model in AstronOS using Cases, Tasks, and Scenario Packs across central and local execution. We compare five complete strategies for carrying an established software-version update plan into a fresh model session: rereading original materials, replaying full history, deterministic text summary, deterministic JSON, and the AstronOS runtime-mediated handoff. Ten controlled tasks are run under all five strategies with three repetitions, yielding 150 included executions. On the single-stage reference family, strategies perform similarly. In the primary three-stage A-C batch, AstronOS passes the frozen scorer in 14 of 15 executions, compared with 0 of 15 for rereading and 2 of 15 for full-history replay; later non-interleaved summary and JSON batches each pass 0 of 15. AstronOS has lower attempt-accounted model-token cost per passing execution, while requiring more execution-window time per attempt. These results associate the complete AstronOS condition with higher end-to-end pass rates across fresh sessions in this benchmark, at a measurable time cost.

Comments23 pages, 2 figures, 13 tables. Zhenhang Nie and Gui Zheng contributed equally; Gui Zheng and Bin Zhang are corresponding authors

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑