发表机构
University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM指令遵循中复制粘贴的捷径问题,提出UNSPECIFIC框架,通过合成共有的约束、强化易满足约束、在文章及其摘要上评估来破解该问题,构建基准并验证其有效性。
AI 中文摘要
大语言模型(LLMs)如今越来越需要遵循复杂指令中的一长串约束,而从参考文档合成指令(即反向翻译)是衡量或增强LLMs遵循复杂指令能力的常用方法。但该方法存在一个关键漏洞:约束合成模型会将参考文本复制为非常具体的约束,被评估的LLM只需在响应中复制该文本就能轻易满足约束。为解决这些问题,我们提出了UNSPECIFIC这一新颖框架,该框架会合成两篇相似参考文章共有的约束以减少复制粘贴,仅选择性强化那些被轻易满足的约束以平衡难度与自然度,并在生成的文章及其摘要上评估约束满足度,以惩罚表面化的指令遵循。我们在新闻、故事和博客领域构建了UNSPECIFIC基准,用于分析LLMs的复制粘贴行为。结果表明,我们合成的约束不仅更具挑战性(例如,GPT-5 Mini的满足率从90%降至78%),且从人类视角看更自然(LLM胜率差距提升30%),还能缓解复制粘贴问题。我们还发现,很大一部分约束被表面化满足(即未在文章核心叙事中得到满足)。代码和数据集已发布在该https URL。
英文摘要
Large language models (LLMs) are increasingly expected to follow long lists of constraints in complex instructions, and synthesizing instructions from a reference document (i.e., back-translation) is a widely used method to measure/enhance LLMs' ability to follow complex instructions. However, this method introduces a critical loophole: the constraint synthesis model copies text from the reference as a very specific constraint and the evaluated LLM trivially satisfies the constraint by copying its text in the response. To address these issues, we propose UNSPECIFIC, a novel framework that synthesizes constraints common to two similar reference articles to reduce copy-pasting, selectively hardens only trivially satisfied constraints to balance difficulty and naturalness, and evaluates satisfaction on both the generated article and its summary to penalize superficial instruction following. Consequently, we built the UNSPECIFIC benchmark on news, story, and blog domains to analyze the copy-pasting behavior of LLMs. Our results show that our synthesized constraints are not only more challenging (e.g., the satisfaction rate of GPT-5 Mini drops from 90% to 78%) and natural (LLM win-rate gap improves by 30%) from a human perspective but also mitigate the copy-pasting. We also find that a large portion of constraints are satisfied superficially (i.e., not satisfied in the core narrative of the article). The code and datasets are released at https://github.com/JeetDSharma/UNSPECIFIC.