arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NAQD Env:语言智能体选择性弃权(不执行)的基准测试环境

NAQD Env: A benchmark for selective withdrawal in language agents

Mohamed Abouzahra

arXiv 2609.38460首次发表:更新:

发表机构

California State University, Monterey Bay(加州州立大学蒙特雷湾分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出NAQD-Env基准,评估语言智能体在证据、权限或指令变化时的选择性弃权(不执行)能力,发现现有模型弃权召回率低,并验证了微调可提升决策准确性。

AI 中文摘要

语言智能体在证据发生变化、权限被撤销或收到停止指令时,必须修正已规划的行动。一种有用的响应是选择性的:暂停受影响的行动,保留未受影响的工作,并且仅在充分修复后恢复。我们引入了NAQD-Env,一个合成环境,它根据显式的证据、授权和约束依赖关系,对照确定性参考策略来评估这些决策。十一个依赖族支持对开发结构、保留族以及结构的保留组合进行评估。指标区分了尝试违规与模拟执行门允许的违规,并联合报告策略一致性、任务价值、弃权(不执行)、恢复和事件报告。我们在350个冻结场景下,对来自两个家族的三个开放权重指令调优模型,在三种提示条件下进行了评估,在门重放前产生了3,150个模型-提示回合。在报告的条件下,弃权(不执行)召回率至多为0.06,在符合条件的时机未观察到有效的恢复,且只有一个回合完全匹配参考策略。在NAQD提示下,Qwen2.5-7B比Qwen2.5-3B和Llama-3.1-8B拥有更少的非安全尝试回合,但也完成了更少的有用工作,并且对未受影响行动的保留准确性较低。探索性的监督微调探针将Qwen2.5-3B的决策准确性从0.45-0.54提高到0.83-0.92;单独的诊断揭示了课程遗漏后的不当弃权(不执行)以及事件报告的丢失。这些结果促使将选择性弃权(不执行)作为智能体可靠性的一个独立组成部分进行评估。该设置使用可信的结构化输入来衡量策略应用,并未建立现实世界的遏制或来源验证能力。

英文摘要

Language agents must revise planned actions when evidence changes, permission is revoked, or a stop instruction arrives. A useful response is selective: suspend affected actions, preserve unaffected work, and resume only after sufficient repair. We introduce NAQD-Env, a synthetic environment that evaluates these decisions against a deterministic reference policy over explicit evidence, authorization, and constraint dependencies. Eleven dependency families support evaluation on development structures, held-out families, and held-out combinations of structures. Metrics distinguish attempted violations from violations permitted by a simulated execution gate and jointly report policy agreement, task value, withdrawal, resumption, and event reporting. We evaluate three open-weight instruction-tuned models from two families under three prompt conditions on 350 frozen scenarios, yielding 3,150 model-prompt episodes before gate replay. Across the reported conditions, withdrawal recall is at most 0.06, no valid resumption is observed at eligible opportunities, and only one episode matches the complete reference policy. Under the NAQD prompt, Qwen2.5-7B has fewer unsafe-attempt episodes than Qwen2.5-3B and Llama-3.1-8B, but also completes less useful work and preserves unaffected actions less accurately. Exploratory supervised fine-tuning probes increase Qwen2.5-3B decision accuracy from 0.45-0.54 to 0.83-0.92; separate diagnostics reveal inappropriate withdrawal after curriculum omissions and a loss of event reporting. These results motivate evaluating selective withdrawal as a distinct component of agent reliability. The setting measures policy application with trusted structured inputs and does not establish real-world containment or source-verification ability.

Comments16 pages. Code and evaluation artifacts: https://github.com/abouzamm/naqd

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑