arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当可验证的计数依赖于措辞:审计指令遵循中的措辞鲁棒性

When Verifiable Counts Depend on Wording: Auditing Wording Robustness in Instruction Following

Qishi Zhan, Seoyeon Jang, Zihan Dong, Minxuan Hu, Ziheng Chen, Tonghui Qu

arXiv 2610.05278首次发表:更新:

发表机构

Marquette University; University of California, San Diego; Georgia Tech; Cornell University; The University of Texas at Austin; Hikvision(马凯特大学; 加利福尼亚大学圣迭戈分校; 佐治亚理工学院; 康奈尔大学; 得克萨斯大学奥斯汀分校; 海康威视)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出WISE评估套件,审计指令遵循模型对措辞变化的鲁棒性,发现仅改变措辞即可显著影响合规性、模型排名及稳定性。

AI 中文摘要

可验证的指令遵循基准通常通过一个固定模板来表达每个约束。我们测试当操作要求不变但措辞变化时,分数是否保持稳定。我们引入了WISE,一个匹配的评估套件和报告协议,实例化于精确单词计数、关键词恰好包含一次以及包含性的8-12单词范围。在100个匹配任务中,来自七个提供商的十三个模型,以及重复生成在完整可见输出上评分,仅措辞就产生了显著的合规性变化。在一个回避族面板中,五种回避和排除形式低于积极基线,而构造性控制也显著改变了合规性:在九模型控制面板中,原始积极形式的合规性为54.9%,较长积极形式为48.2%,目标出现较晚时为36.7%,AVOID1为33.8%。一个严格的JSON结构探针显示了超出计数的措辞敏感性,且影响方向不同。效应大小、失败方向、最弱形式和模型排名因实现而异。在最具破坏性的排除形式下,排名最高的模型发生变化,24.1%的严格排序模型对发生逆转。人工验证进一步表明,对精确计数解释的一致同意可以与显著不同的模型行为共存。WISE用平均和最差形式合规性、措辞差距、失败概况和排名稳定性补充了传统分数。

英文摘要

Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword inclusion exactly once, and an inclusive 8--12 word range. Across 100 matched tasks, up to thirteen models from seven providers, and repeated generations scored over the complete visible output, wording alone produces substantial compliance shifts. In an avoidance-family panel, five avoidance and exclusion forms fall below the positive baseline, while constructional controls also shift compliance substantially: in the nine-model control panel, compliance is 54.9% for the original positive form, 48.2% for a longer positive form, 36.7% when the target appears later, and 33.8% for AVOID1. A strict JSON-structure probe shows wording sensitivity beyond counting, with a different direction of effect. Effect sizes, failure directions, weakest forms, and model rankings vary across realizations. Under the most disruptive exclusion form, the top-ranked model changes and 24.1% of strictly ordered model pairs reverse. Human validation further shows that unanimous agreement on an exact-count interpretation can coexist with substantially different model behavior. WISE supplements conventional scores with mean and worst-form compliance, wording gaps, failure profiles, and ranking stability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑