arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自由生成分支,审慎行动:面向递归大语言模型智能体树的渐进式风险归属机制

Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees

Molly Wang

arXiv 2609.01035首次发表:更新:

发表机构

Imperial Business School(帝国商学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对递归大语言模型智能体树,提出渐进式风险归属机制以管控分支行动权限,证明其危害边界,通过模型与实验验证规则,为智能体安全设计提供依据。

AI 中文摘要

递归大语言模型(LLM)智能体可通过生成专门分支来拓宽搜索范围,部分分支后续会请求发送数据或部署代码的工具。问题在于,分支何时应获得行动权限?我们区分两种机制:沙箱生成(即外部控制可避免特定危害)与能力激活(即选定分支跨越不可逆行动边界)。渐进式风险归属(PRV)机制会将轨迹级风险预算托管,并在分支激活时扣除该预算。我们证明了自适应生成树的任意时刻危害边界。分支结果可能存在依赖关系,但每个局部证明需在激活前完整历史(包括用于选择请求的信息)条件下保持有效。当激活门、分支扣费和计算约束固定时,延迟归属机制保留了不可撤销分支扣费下的所有策略。分支选择后,边际风险估计仍可能失效。在一个简化分支模型中,轨迹危害随权威再生数$\boldsymbol{\textit{R}}_A$跨越1而变化:当局部风险$p$趋近于0时,轨迹危害在临界值以下与$p$成正比,在临界值处与$\boldsymbol{\textit{p}}$的平方根成正比,且在临界值以上保持正下限。有限类型占用模型可得出风险与计算影子价格。对于单位风险边际价值递减的嵌套扇出模式,这些价格产生阈值规则。分支计算与拆分样本实验验证了结果,但这些合成研究未评估部署智能体的安全性。分析提出设计规则:在沙箱中广泛搜索,并以明确风险扣费审慎授予递归权限。

英文摘要

Recursive LLM agents can broaden their search by spawning specialists. Some branches later request tools that send data or deploy code. When should a branch receive authority to act? We distinguish sandbox spawning, in which external controls prevent the specified harm, from capability activation, in which a selected branch crosses an irreversible-action boundary. Progressive Risk Vesting (PRV) holds a trajectory-level risk budget in escrow and debits it as branches are activated. We prove an anytime harm bound for adaptively generated trees. Branch outcomes may be dependent, but each local certificate needs to remain valid conditional on the full pre-activation history, including the information used to select the request. When activation gates, branch charges, and compute constraints are held fixed, delayed vesting preserves every policy available under irrevocable spawn charging. Marginal risk estimates can still fail after branch selection. In a stylized branching model, trajectory harm changes as the authority reproduction number $\mathcal{R}_A$ crosses one. As local risk $p$ approaches zero, trajectory harm is proportional to $p$ below criticality, proportional to $\sqrt{p}$ at criticality, and retains a positive floor above it. A finite-type occupancy model yields risk and compute shadow prices. For nested fanout modes with decreasing marginal value per unit risk, these prices produce a threshold rule. Branching calculations and a split-sample experiment illustrate the results. These synthetic studies do not estimate safety in deployed agents. The analysis suggests a design rule: search broadly in the sandbox and grant recursive authority sparingly, with an explicit risk charge.

Comments8 pages, 2 figures. Theory and synthetic numerical studies; no deployed-agent evaluation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑