arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11727cs.AI

Harness-IF:评估编码智能体在不同指令层面的指令遵循能力

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

  • Tsinghua University(清华大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

Zining Huang, Haoran Que, Hong Zeng, Ge Zhang, Zuo Wang, Jin Chen, Haodong Wang, Zhongfei Hou, Changxin Pu, Shen Yan, Wenhao Huang

AI总结:

本研究推出Harness-IF评估框架,引入AP-Acc指标区分编码智能体的指令合规与巧合,发现12个前沿模型在对抗先验规则上表现更差,且合并优先级不随提示深度变化。

AI中文摘要:

当编码智能体遵守某条规则时,它可能本来就会这么做,现有指令遵循基准无法区分这种情况:它们将规则集中在用户轮次,而编码智能体基准则强调最终任务的成功。我们推出Harness-IF,该方法从执行证据中逐一对操作规则进行评分:从包含642条规则的库中抽取60个真实的多轮编码条目,对256条规则进行判定,并将其置于已部署智能体读取的5个可配置层面上。为了区分合规性与巧合,我们引入了先验对抗准确率(AP-Acc),该指标仅对标记为与未提示默认值相反的规则进行评分,方法是在9个探测构建中省略该规则,其余部分保持不变,重新运行任务。在12个前沿模型中,准确率介于72.1%-85.9%之间,AP-Acc介于66.1%-78.6%之间;所有模型在对抗先验规则上的表现都更差,差距为3.6至7.4个百分点(平均5.81),且该方向在具有条目聚类区间的共同支持分析中依然存在。因此,聚合评分以特定于模型的幅度高估了合规性:先验控制使排名靠前的构建保持不变,并交换了三对相邻排名。在9个独立构建上进行的平衡冲突试点增加了第二个结果:合并优先级不遵循提示深度,系统提示、项目文件和用户指令的优先级高于工具和技能描述。

英文摘要:

When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate rules in the user turn, while coding-agent benchmarks emphasize final task success. We introduce Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads. To separate compliance from coincidence we introduce Against-Prior Accuracy (AP-Acc), which scores only rules labeled as opposing unprompted defaults, observed by re-running tasks with the rule withheld across nine probe builds and curated otherwise. Across 12 frontier models, accuracy spans 72.1-85.9% and AP-Acc 66.1-78.6%; every model is worse on against-prior rules, by 3.6 to 7.4 points (mean 5.81), and the direction survives a common-support analysis with item-clustered intervals. Aggregate scores therefore overstate compliance by a model-specific margin: prior control leaves the top build unchanged and exchanges three adjacent rank pairs. A counterbalanced conflict pilot on nine separate builds adds a second result: pooled precedence does not follow prompt depth, with system prompts, project files, and user instructions ahead of tool and skill descriptions.

补充信息

↑