发表机构
McGill University; Polytechnique de Montréal(麦吉尔大学; 蒙特利尔综合理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
G-CARB通过沿依赖关系选择证据并利用共形风险预算控制智能体停止时机,在降低监控开销的同时提升任务完成率。
AI 中文摘要
小型语言模型(SLM)智能体需要安全控制,以便在监控开销极低的情况下跟踪跨工具调用的后果。例如,一次私有读取在后续操作将该数据发送到系统外部时,就变成了泄露。我们引入了CARB(共形智能体风险预算),它利用停止前已产生的危害台账来校准何时停止智能体。在可交换情节下,标准共形风险控制可在校准和未来情节的期望上限制这一声明损失。G-CARB沿从私有源到外部操作的可见依赖关系选择评分器证据。台账仍覆盖整个已执行历史,且计算门控分数无需额外的语言模型推理。在AgentDojo回放中,使用两个14B骨干模型,G-CARB在中等风险预算下将评分器输入记录大致减半,同时相对于全前缀评分提高了自主任务完成率;相同大小的随机上下文也获得了类似的增益。受控示例表明,保留相关依赖可进一步避免中断良性工作。
英文摘要
Small language model (SLM) agents need safety controls that track consequences across tool calls with little monitoring overhead. A private read, for example, becomes a leak when a later action sends that data outside the system. We introduce CARB (Conformal Agent Risk Budget), which calibrates when to stop an agent using a ledger of harm incurred before stopping. Under exchangeable episodes, standard conformal risk control bounds this declared loss in expectation over calibration and a future episode. G-CARB selects scorer evidence along observable dependencies from private sources to outgoing actions. The ledger still covers the entire executed history, and computing the gate score requires no additional language-model inference. On AgentDojo replay with two 14B backbones, G-CARB roughly halves scorer-input records at intermediate risk budgets while improving autonomous task completion relative to full-prefix scoring; random context of the same size achieves similar gains. Controlled examples show how retaining the relevant dependency can further avoid stopping benign work.
CommentsSLMs for Agentic Systems, Paris, France, 2026