arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向AI安全评估的对抗语用学:指令冲突、嵌入命令与策略模糊性基准

Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control

Brett Reynolds

arXiv 2607.01153首次发表:更新:

发表机构

Humber Polytechnic; University of Toronto(汉博理工学院; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出对抗语用学基准和标注协议,通过语言学控制的分类法评估模型在指令冲突、嵌入命令等场景下的行为,为安全评估提供实证和方法论工具。

AI 中文摘要

语言模型的安全评估越来越依赖于对模糊自然语言行为的判断:模型是否遵循了指令、是否恰当拒绝、是否遵守策略、是否抵抗嵌入命令、或在代理任务中错误报告进展。现有基准通常将这些区别压缩为通过/失败标签,掩盖了失败是源于能力限制、策略模糊性、指令冲突、支架故障还是评估者判断不稳定。本文引入对抗语用学作为基准和标注协议,用于评估模型在指令冲突、嵌入命令、引用、范围模糊性、指示语、间接言语行为和多轮代理转录本下的行为。贡献是经验性和方法论的:一个语言学控制的分类法、一个带有验证者强制元数据的18项种子基准、一个54行本地种子试点、一个区分任务成功、策略合规、安全风险、拒绝结果和评估者置信度的专家评估协议,以及用于判断有效性、诊断模糊性和分类法漂移的指标。该框架将语言判断方法论转化为验证安全评估、LLM判断器、金标准构建、提示注入测试和安全文档的实用工具。

英文摘要

Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately, complied with a policy, or misreported progress in an agentic task. Existing benchmarks compress these into pass/fail labels, obscuring whether failures reflect capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. Adversarial pragmatics is safety-relevant model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, and indirect speech acts. It's designed to extend to multi-turn agent transcripts, but the seed set represents that family with a single-turn tool-result contrast. This paper introduces a diagnostic framework, an 18-item seed benchmark, a 54-row pilot, and a six-cell LLM-judge assessment, with a protocol keeping task success, policy compliance, risk, refusal, attribution, and confidence analytically separate. The benchmark separates four inference targets a single label can conceal: the regime-relative reference, configured-system behaviour, evaluator-output interpretation, and taxonomic assignment. Its intended use is diagnosis, not deployment certification, vendor ranking, or a general safety score. A first LLM judge that graded its own outputs with the expected answer visible missed the safety-relevant minority classes. Item-clustered intervals leave four of six chance-corrected agreement statistics unable to rule out a constant labeller, and hierarchical pooling shrinks the one eye-catching rubric effect toward the group mean and widens its interval through zero. Rejudging the objects across three judge models and two information conditions leaves the pattern intact: no cell recovers more than two of eleven partial successes, and the strongest cell's edge comes partly from never using that label.

Comments32-page main paper plus 13-page supplement; 6 figures and 17 tables total; code and data artifact available at the linked repository

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑