\textsc{IH - 基准测试}:用于大语言模型应用中指令层次鲁棒性的以冲突为中心的基准测试
IH-Benchmark: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications
浏览论文内容
中文总结 AI 辅助
研究大语言模型接收到不同优先级冲突指令时遵循哪个的问题,提出IH - 基准测试,基于人工分类法构建,用统一协议评估,涵盖多种设置。通过对37个模型评估发现指令层次鲁棒性非单一能力,需多方面评估。
中文摘要 AI 辅助
当语言模型接收到来自不同优先级的冲突指令时,它实际遵循哪一个?这个问题是可靠的大语言模型部署的核心。现有基准测试只能部分回答这个问题,通常只关注单个层次边缘或改编工具使用覆盖有限的公共数据集。我们提出了IH - 基准测试,这是一个以冲突为中心的基准测试,用于跨直接系统 - 用户冲突(S>U)和工具介导的用户 - 工具(U>T)冲突的指令层次鲁棒性。IH - 基准测试基于44个约束家族的人工分类法构建,涵盖通用、健康、金融、零售和编码设置,并使用统一的二进制通过/失败协议评估场景,该协议将谓词DSL与类别范围的大语言模型判断相结合。在37个评估模型中,层次合规率从98.2%到20.5%不等。我们发现,强大的S>U合规性不是U>T鲁棒性的可靠指标:一些模型在直接用户冲突下保留系统约束,但当工具输出中出现冲突指令时会急剧下降。约束强化还揭示了模型之间的差异:一些故障在更强的警告下基本得到修复,而另一些故障在所有严格级别上都持续存在。最后,最具揭示性的故障往往很微妙而不是明显危险;模型抵制未经授权的购买或批量票务关闭比注入免责声明或小事实扭曲更可靠。这些结果表明,指令层次鲁棒性不是单一能力,而是一组必须跨冲突表面、约束类型和攻击呈现进行评估的行为。
英文摘要
When a language model receives conflicting instructions from different priority levels, which one does it actually follow? This question lies at the heart of reliable LLM deployment. Existing benchmarks answer this only partially, often focusing on a single hierarchy edge or adapting public datasets with limited tool-use coverage. We present IH-Benchmark, a conflict-centered benchmark for instruction-hierarchy robustness across direct system-user conflicts (S>U) and tool-mediated user-tool (U>T) conflicts. IH-Benchmark is built from a human-authored taxonomy of 44 constraint families across generic, health, finance, retail, and coding settings, and evaluates scenarios with a uniform binary pass/fail protocol combining a predicate DSL with category-scoped LLM judges. Across 37 evaluated models, hierarchy compliance ranges from 98.2% to 20.5%. We find that strong S>U compliance is not a reliable proxy for U>T robustness: several models preserve system constraints under direct user conflict but degrade sharply when conflicting instructions appear in tool outputs. Constraint hardening also reveals a split between models: some failures are largely fixed by stronger warnings, while others persist across all strictness levels. Finally, the most revealing failures are often subtle rather than overtly dangerous; models resist unauthorized purchases or bulk ticket closure more reliably than injected disclaimers or small factual distortions. These results suggest that instruction-hierarchy robustness is not a single capability, but a set of behaviors that must be evaluated across conflict surfaces, constraint types, and attack presentations.