arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

G-CARB:面向组合危害的图局部化共形智能体风险预算

G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm

Zijun Yu, Yu Gu, Vahid Partovi Nia, Masoud Asgharian

arXiv 2610.05563首次发表:更新:

发表机构

McGill University; Polytechnique de Montréal(麦吉尔大学; 蒙特利尔综合理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

G-CARB通过沿依赖关系选择证据并利用共形风险预算控制智能体停止时机,在降低监控开销的同时提升任务完成率。

AI 中文摘要

小型语言模型(SLM)智能体需要安全控制,以便在监控开销极低的情况下跟踪跨工具调用的后果。例如,一次私有读取在后续操作将该数据发送到系统外部时,就变成了泄露。我们引入了CARB(共形智能体风险预算),它利用停止前已产生的危害台账来校准何时停止智能体。在可交换情节下,标准共形风险控制可在校准和未来情节的期望上限制这一声明损失。G-CARB沿从私有源到外部操作的可见依赖关系选择评分器证据。台账仍覆盖整个已执行历史,且计算门控分数无需额外的语言模型推理。在AgentDojo回放中,使用两个14B骨干模型,G-CARB在中等风险预算下将评分器输入记录大致减半,同时相对于全前缀评分提高了自主任务完成率;相同大小的随机上下文也获得了类似的增益。受控示例表明,保留相关依赖可进一步避免中断良性工作。

英文摘要

Small language model (SLM) agents need safety controls that track consequences across tool calls with little monitoring overhead. A private read, for example, becomes a leak when a later action sends that data outside the system. We introduce CARB (Conformal Agent Risk Budget), which calibrates when to stop an agent using a ledger of harm incurred before stopping. Under exchangeable episodes, standard conformal risk control bounds this declared loss in expectation over calibration and a future episode. G-CARB selects scorer evidence along observable dependencies from private sources to outgoing actions. The ledger still covers the entire executed history, and computing the gate score requires no additional language-model inference. On AgentDojo replay with two 14B backbones, G-CARB roughly halves scorer-input records at intermediate risk budgets while improving autonomous task completion relative to full-prefix scoring; random context of the same size achieves similar gains. Controlled examples show how retaining the relevant dependency can further avoid stopping benign work.

CommentsSLMs for Agentic Systems, Paris, France, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑