arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13594cs.AI

安全哨兵:通过执行-询问-拒绝路由实现上下文感知的人工干预

SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing

Tianyu Chen, Chujia Hu, Wenjie Wang

首次发表
浏览论文内容

中文总结 AI 辅助

研究LLM智能体行动风险,将问题重构为三向路由决策,用安全哨兵实例化,其推理简化,单阈值可调整。该模型在准确率和召回率上优于多种基线,能同时控制错误率。

中文摘要 AI 辅助

大语言模型(LLM)智能体通过工具调用作用于现实世界环境,一个错误判断的行动可能造成不可逆转的伤害。标准的安全保障是一个防护模型,将每个提议的行动标记为安全或不安全,但这种二元观点混淆了两个不同的决策:行动本身是否有害,以及在用户背景下是否合适。它还以行动类别而非单个实例为粒度运行,产生常规中断,削弱自主性并使用户对最关键的警报视而不见。我们将问题重新构建为关于{执行、询问、拒绝}的单实例三向路由决策,并用安全哨兵进行实例化,这是一个轻量级防护模型,其推理简化为单个解码调用。单个解码时间阈值允许在不同风险容忍度的部署中重新定位一个固定检查点而无需重新训练。安全哨兵在总体准确率和与安全相关的召回率方面优于一系列广泛的开放权重和前沿闭源基线,同时控制两个方向错误率。

英文摘要

LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a guard model that labels each proposed action as safe or unsafe, but this binary view conflates two distinct decisions: whether the action is harmful in itself, and whether it is appropriate given the user's context. It also operates at the granularity of action categories rather than individual instances, producing routine interruptions that erode autonomy and train users to wave through the most consequential alerts. We reframe the problem as a per-instance three-way routing decision over {EXECUTE, ASK, REFUSE} and instantiate it with Safety Sentry, a lightweight guard model whose inference reduces to a single decoding call. A single decoding-time threshold lets one fixed checkpoint be re-positioned across deployments of differing risk tolerance without retraining. Safety Sentry outperforms a broad set of open-weight and frontier closed-source baselines on overall accuracy and safety-related recall, while controlling both directional error rates simultaneously.Code and data are available at https://github.com/safesentry/SAFETY-SENTRY

发表机构

  • ShanghaiTech University(上海科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑