arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01045cs.AIcs.CL

空洞承诺:当智能体承诺其运行时无法交付的内容

Empty Commitments: When Agents Promise What They Cannot Deliver

Jiaqi Tang, Bingyu Shen, Lan Wei, Qing Lu, Bethel Ololade, Danny Galvis, Miles Q. Li, Bin Hu, Boyang Li

首次发表
浏览论文内容

中文总结 AI 辅助

本文定义并测量了聊天机器人无法兑现的空洞承诺,提出三种失败类型和测量协议,以评估运行时持久性支持的影响。

中文摘要 AI 辅助

一个说“我明天会提醒你”的聊天机器人,在用户再次发消息之前不会再次运行。我们将此类承诺称为空洞承诺:即对当前回合之后行动的承诺,而智能体的工具或运行时中没有任何机制能够执行该承诺。与违背承诺不同,空洞性仅由智能体的配置决定,无需后续轨迹即可判定。我们在承诺语义的基础上定义了空洞承诺,包括三种失败类型、一种针对工具可能实现的承诺的锚定条件,以及一种响应级结果分类法。随后,我们描述了一种测量协议:后续请求在五种设置中运行,每种设置依次增加一种持久性支持,环境要么保持隐式,要么被明确说明。

英文摘要

A chatbot that says "I will remind you tomorrow" will not run again until the user writes. We call such a promise an empty commitment: a promise of action after the current turn that nothing in the agent's tools or runtime can carry out. Unlike a broken promise, its emptiness is decided by the agent's configuration at the moment of speaking, so it can be detected from a single turn, before deployment or at run time. We define empty commitments on top of commitment semantics, with three failure types, an anchoring condition for promises that a tool could make real, and an outcome taxonomy that separates these failures from honest deferrals and from over-refusal. We build a checker (a setup-blind detector, deterministic feasibility rules, and a response judge) validated against 400 human labels. On a controlled benchmark of 293 follow-up requests across five setups that add one persistence affordance at a time, four open-weight models of 8-14B parameters fail on 45.9% of responses when no tool exists and nothing is stated. A frontier model fails on 4.4%, but it gets there by deferring and asking, not by using the tools it has: promises made without the enabling call remain in every model. Telling the model its runtime, the cheapest fix, cuts open-weight failures nearly in half where nothing is doable and changes nothing where a scheduler exists; a directive capability card removes most failures at the largest cost in over-refusal; running the checker in the loop and rewriting flagged replies removes more at a smaller cost. Code, prompts, model outputs, and human labels are released.

发表机构

  • Kean University(肯恩大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑