arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SafeCommit:验证基于记忆的智能体何时可以安全行动

SafeCommit: Certifying When Memory-Grounded Agents May Safely Act

Mayur Akewar, Ravi Ranjan

arXiv 2608.04289首次发表:更新:

发表机构

Florida International University(佛罗里达国际大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长程智能体过早承诺的问题,提出SafeCommit风险控制层,通过构建合理潜在世界、共形动作证书等实现安全行动决策,还展示了安全-效用权衡。

AI 中文摘要

长程智能体越来越多地使用持久记忆和工具来采取具有外部副作用的行动。一个核心失败模式是过早承诺:智能体在未解决其记忆基础是否过时、冲突、不完整或损坏的情况下就采取行动。我们将此问题形式化为记忆不确定性下的安全承诺,并引入SafeCommit,这是一个介于智能体推理和外部执行之间的风险控制层。该层从记忆、观测、工具输出、来源和策略约束中构建一组校准的合理潜在世界。仅当共形动作证书表明该动作在每个保留的世界中都是安全的时,它才允许具有副作用的动作;否则,它会选择针对阻碍认证的世界的低副作用探测,或返回保守的 fallback。在校准的世界覆盖率下,不安全的已认证承诺的概率最多为目标水平α;在世界提议不完善的情况下,该界限区分了校准误差和表示误差。一个无依赖的受控模拟器展示了安全-效用权衡,并通过一条命令复现所有报告的结果。目标是提供一种具体方法,不仅决定智能体应做什么,还决定何时现有证据足以安全地做这件事。

英文摘要

Long-horizon agents increasingly use persistent memory and tools to take actions with external side effects. A central failure mode is premature commitment: an agent acts before resolving whether its memory grounding is stale, conflicting, incomplete, or corrupted. We formalize this problem as safe commitment under memory uncertainty and introduce SafeCommit, a risk controlled layer between agent reasoning and external execution. The layer constructs a calibrated set of plausible latent worlds from memory, observations, tool outputs, provenance, and policy constraints. It permits a side effectful action only when a conformal action certificate shows that the action is safe in every retained world. Otherwise, it selects a low-side-effect probe that targets the worlds blocking certification, or returns a conservative fallback. Under calibrated world coverage, the probability of an unsafe certified commit is at most the target level α; with imperfect world proposal, the bound separates calibration and representation error. A dependency-free controlled simulator illustrates the safety-utility tradeoff and reproduces all reported results with one command. The goal is to offer a concrete approach for deciding not only what an agent should do, but when the available evidence is sufficient to safely do it.

Comments14 pages, 6 tables, and 1 figure, target NeurIPS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑