arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentR:面向基于大语言模型的可审计工作流的有状态且故障恢复感知的软件架构

AgentR A Stateful and Recovery-Aware Software Architecture for LLM-based Auditable Workflows

Riya Samanta, Bidyut Saha, Soumya Kanti Ghosh, Rajkumar Buyya

arXiv 2608.15264首次发表:更新:

AI 中文总结

提出AgentR架构,以有状态设计结合异步编排与成本审计,实现LLM工作流的高可恢复性、可观测性与可审计性,原型在文献综述场景验证获99.2%作业完成率及最高4.3倍并行加速

AI 中文摘要

现代基于大语言模型(LLM)的应用日益需要多阶段执行、持久化中间状态、重试语义以及可审计的使用记账。然而,许多LLM应用仍被实现为无状态的提示-响应包装器或会话受限的对话系统,这使得它们在中断或故障后难以恢复、审计和复现。我们提出AgentR,一种用于LLM工作流系统的有状态架构,该架构支持持久化和故障恢复,并以科学文献综述作为代表性用例实例化。AgentR将研究意图、生成的查询、候选论文评估以及差距分析表示为持久化工作流工件,并通过由Redis支持的异步BullMQ工作者执行管道,以PostgreSQL作为持久化存储。该设计包含显式处理状态转换、指数退避重试、孤立作业检测、信用感知预检查、ACID代币成本日志以及Type 2缓慢变化的定价记录。我们基于原型部署收集的遥测数据评估AgentR。在LLM阶段,系统实现99.2%的作业完成率,且意图分解、查询生成和论文评分的平均延迟分别为9.0秒、18.9秒和25.4秒。并行评分允许基于观测调用进行分析延迟建模,相较于顺序执行可实现高达4.3倍的墙钟速度提升。这些结果提供了初步的概念验证,表明持久化状态机设计、异步编排以及成本感知的使用日志可提升LLM工作流系统的可观测性、可恢复性和运营问责性。AgentR的原型实现可公开获取:https://github.com/RiyaSamanta/AgentR-public

英文摘要

Modern LLM-based applications increasingly require multi- stage execution, persistent intermediate state, retry seman- tics, and auditable usage accounting. However, many LLM applications are still implemented as stateless prompt- response wrappers or session-bounded conversational sys- tems, which makes them difficult to recover, audit, and re- produce after interruption or failure. We propose AgentR, a stateful architecture for LLM workflow systems that en- ables persistence and recovery, instantiated through scien- tific literature review as a representative use case. AgentR represents research intent, generated queries, candidate- paper assessments and gap analyses as durable workflow artifacts, and executes the pipeline through asynchronous BullMQ workers backed by Redis, with PostgreSQL as the persistence store. The design includes explicit processing state transitions, retries with exponential backoff, orphan job detection, credit-aware pre-checks, ACID token-cost logging, and Type-2 slowly changing pricing records. We evaluate AgentR on telemetry collected from a prototype deployment. At the LLM stage, the system achieves 99.2% job completion, and mean latencies of 9.0 s, 18.9 s, and 25.4 s for intent decomposition, query generation, and paper scoring, respectively. Parallel scoring allows for analytical latency modeling from observed calls, leading to as much as 4.3 wall-clock speedup over sequential execution. The results provide preliminary proof-of-concept that persistent state machine design, asynchronous orchestration, and cost-aware usage logging can enable improved observability, recoverability, and operational accountability in LLM workflow systems. The prototype implementation of AgentR is publicly available at: https://github.com/ RiyaSamanta/AgentR-public.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑