AI 中文总结
该研究针对自主智能体对传统LLM服务的挑战,提出Aries全栈实验框架,通过实验揭示服务系统的关键瓶颈,探讨智能体原生服务系统的设计愿景。
AI 中文摘要
自主智能体通过将重复推理与持久上下文、沙盒化工具执行相结合,对传统大语言模型(LLM)服务提出了挑战。我们提出Aries,这是一个全栈实验框架,它将任务语义与执行配置分离,通过关联的系统遥测数据重构跨组件智能体轨迹,并在异构沙盒 substrate 上通过一致接口暴露有状态工具执行。我们使用Aries在开源智能体 harness 和基准上进行可复现实验,并补充了来自商业平台的生产追踪数据,将底层系统研究建立在观测到的生产行为基础上。结果显示:(1)以token为中心的指标会遗漏非推理瓶颈;(2)保留额外上下文会带来递减的准确性收益,同时降低服务容量;(3)工具沙盒在长空闲期与短资源突发期之间交替,而当前基于快照的状态管理使得激进弃权(不执行)的成本高昂。补充安全分析进一步强调了减少沙盒攻击面的必要性。随后我们讨论围绕轨迹级指标、自适应上下文管理、弹性沙盒资源管理及最小化攻击面的沙盒设计的智能体原生服务系统愿景。
英文摘要
Autonomous agents challenge conventional LLM serving by coupling repeated inference with persistent context and sandboxed tool execution. We present Aries, a full-stack experimentation framework that separates task semantics from execution configurations, reconstructs cross-component agent trajectories with correlated system telemetry, and exposes stateful tool execution through a consistent interface across heterogeneous sandbox substrates. We use Aries to conduct reproducible experiments on open agent harnesses and benchmarks. We complement these experiments with production traces from a commercial platform, grounding low-level systems research in observed production behavior. Our results show that (1) token-centric metrics miss non-inference bottlenecks, (2) retaining additional context yields diminishing accuracy benefits while reducing serving capacity, and (3) tool sandboxes alternate between long idle periods and short resource bursts, while current snapshot-based state management makes aggressive suspension costly. A complementary security analysis further highlights the need to reduce the sandbox attack surface. We then discuss the vision for agent-native serving systems designed around trajectory-level metrics, adaptive context management, elastic sandbox resource management, and sandboxes with minimized attack surface.