解构指令遵循:一个用于大语言模型指令合规能力细粒度评估的新基准
Deconstructing Instruction-Following: A New Benchmark for Granular Evaluation of Large Language Model Instruction Compliance Abilities
- Card Intelligence, Capital One(Card Intelligence,Capital One)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出MOSAIC基准,通过细粒度分析揭示大语言模型在复杂指令遵循中的差异与弱点,为提升模型可靠性提供关键洞察。
AI中文摘要:
可靠地确保大型语言模型(LLMs)遵循复杂指令是一个关键挑战,因为现有基准往往无法反映现实世界的应用或隔离合规性与任务成功。我们介绍了MOSAIC(MOdular Synthetic Assessment of Instruction Compliance),一个模块化框架,它使用动态生成的数据集,包含多达20个应用导向的生成约束,以实现对这一能力的细粒度和独立分析。基于此新基准对五个不同家族的LLM进行评估,证明合规性并非单一能力,而是显著受到约束类型、数量和位置的影响。分析揭示了模型特定的弱点,揭示了指令之间的协同与冲突相互作用,并识别了不同的位置偏差,如优先效应和近期效应。这些细粒度的见解对于诊断模型故障和开发更可靠的LLM至关重要,以满足对复杂指令严格遵守的需求。
英文摘要:
Reliably ensuring Large Language Models (LLMs) follow complex instructions is a critical challenge, as existing benchmarks often fail to reflect real-world use or isolate compliance from task success. We introduce MOSAIC (MOdular Synthetic Assessment of Instruction Compliance), a modular framework that uses a dynamically generated dataset with up to 20 application-oriented generation constraints to enable a granular and independent analysis of this capability. Our evaluation of five LLMs from different families based on this new benchmark demonstrates that compliance is not a monolithic capability but varies significantly with constraint type, quantity, and position. The analysis reveals model-specific weaknesses, uncovers synergistic and conflicting interactions between instructions, and identifies distinct positional biases such as primacy and recency effects. These granular insights are critical for diagnosing model failures and developing more reliable LLMs for systems that demand strict adherence to complex instructions.