arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于智能体编排的形式化分层架构,基于栈的执行和惰性发现

A Formal Hierarchical Architecture for Agentic Orchestration with Stack-Based Execution and Lazy Discovery

Prashant Devadiga, Abhishek, Adithya Mishra, Alok Singh, Amisha Sinha, Asit Desai, Gaurang Dahad, Harshit Bhushan, Mandati Pramod Reddy, Prakhar Gupta, Rupesh Patil, Siddhi Behere

arXiv 2607.11138首次发表:更新:

AI 中文总结

研究大语言模型智能体架构瓶颈,提出分层基于技能的编排架构,通过LIFO栈和惰性发现协议解决问题,防止输出泄漏,并进行数学形式化、算法分析及基准测试,为企业环境部署提供隔离保证。

AI 中文摘要

大语言模型智能体能力的快速扩展暴露出一个关键的架构瓶颈:当智能体访问扁平的整体工具注册表时,模型必须同时评估数百或数千个选项,导致决策空间爆炸、上下文窗口饱和以及路由准确性下降。为解决这些限制,本文提出一种分层的、基于技能的智能体编排架构。能力被组织成一棵有根树,内部节点进行路由决策,叶节点执行确定性任务。运行时通过后进先出(LIFO)栈强制单步执行循环,赋予智能体类似下推自动机的内存形式,使其能跟踪嵌套执行上下文并从任何深度确定性恢复。能力发现遵循清单驱动的惰性加载协议,仅加载活动节点的直接子节点,因此内存和提示成本随探索路径而非全局注册表缩放。通过用局部栈帧替换全局内存,该架构防止一个执行分支的输出泄漏到另一个分支,为在受监管的企业环境中部署建立所需的隔离保证。我们还讨论了UPI Help,一个由人工智能驱动的数字支付支持产品,作为一个激励性的生产部署场景。我们提供了编排状态的数学形式化、执行循环的详细算法分析,以及在不断增加的工具目录、多步工作流压力和每个大语言模型调用的可见模式令牌暴露下比较扁平路由和分层路由的受控基准测试。

英文摘要

The rapid expansion of capabilities in Large Language Model (LLM) agents has exposed a critical architectural bottleneck: when agents are given access to a flat, monolithic registry of tools, the model must evaluate hundreds or thousands of options simultaneously. This leads to decision-space explosion, context window saturation, and degraded routing accuracy. To address these limitations, this paper presents a hierarchical, skill-based architecture for agentic orchestration. Capabilities are organized as a rooted tree where internal nodes make routing decisions and leaf nodes execute deterministic tasks. The runtime enforces a single-step execution loop governed by a Last-In-First-Out (LIFO) stack, giving the agent a form of memory akin to a Pushdown Automaton, therefore enabling it to track nested execution contexts and resume deterministically from any depth. Capability discovery follows a manifest-driven, lazy-loading protocol: only the immediate children of the active node are loaded, so memory and prompt costs scale with the explored path rather than the global registry. By replacing global memory with localized stack frames, the architecture prevents outputs from one execution branch from leaking into another, establishing the isolation guarantees required for deployment in regulated enterprise environments. We also discuss UPI Help, an AI-powered digital payments support product, as a motivating production deployment context. We provide a mathematical formalization of the orchestration state, detailed algorithmic analysis of the execution loop, and controlled benchmarks comparing flat and hierarchical routing under increasing tool catalogs, multi-step workflow pressure, and visible schema-token exposure per LLM call.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑