arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IFHierBench:面向大语言模型的分层指令遵循基准

IFHierBench: Hierarchical Instruction Following for Large Language Models

Yuetian Mao, Chunyang Chen

arXiv 2607.27912首次发表:更新:

发表机构

Technical University of Munich(慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出分层指令遵循基准IFHierBench,评估发现现有大语言模型遵循嵌套约束的能力不足,为提升指令遵循能力的训练方法提供了研究方向。

AI 中文摘要

指令遵循能力对于将大语言模型部署到实际应用中至关重要,下游组件依赖于满足特定约束的输出。现代部署越来越倾向于在单次大语言模型调用中处理完整任务,其中单个提示指定分层输出,其整体产物、结构部分和嵌套字段均需满足具体约束。现有指令遵循基准将约束集视为统一应用于响应的扁平列表,因此无法将检查范围限定到输出的特定部分。我们引入IFHierBench,这是一个包含600个提示的分层指令遵循基准,这些提示按四个约束树深度和35种不同约束分层,每个提示都配有确定性检查器,可在每个范围内验证约束是否得到满足。通过评估七个领先的专有和开放权重模型,我们发现即使是最强的模型,其提示级准确率也仅略超50%,且准确率会随着约束深度的增加而急剧下降。可靠遵循嵌套约束仍是当前大语言模型的一个重大缺口,这为未来的训练方法提供了动机,这些方法需要在更精细的粒度上考虑约束遵循,以实现更好的指令遵循能力。

英文摘要

Instruction-following ability is critical for deploying large language models in real-world applications, where downstream components depend on the output satisfying specific constraints. Modern deployments increasingly handle the full task in a single LLM call, with one prompt specifying a layered output whose overall artifact, structural sections, and nested fields must each satisfy concrete constraints. Existing instruction-following benchmarks treat the constraint set as a flat list applied uniformly to the response, so they cannot scope a check to a particular section of the output. We introduce IFHierBench, a hierarchical instruction-following benchmark of 600 prompts stratified across four constraint-tree depths and 35 distinct constraints, each prompt paired with a deterministic checker that verifies satisfaction at every scope. Evaluating seven leading proprietary and open-weight models, we find that even the strongest model only marginally exceeds 50% prompt-level accuracy and that accuracy degrades sharply as constraint depth grows. Reliably following nested constraints remains a substantial gap for current LLMs, motivating future training methods that consider constraint adherence at finer granularity to achieve better instruction-following ability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑