发表机构
Center for Information and Language Processing, LMU Munich; Munich Center for Machine Learning (MCML); R&D Centre for Large Language Models, National Institute of Informatics, Japan(慕尼黑大学信息与语言处理中心; 慕尼黑机器学习中心; 日本国立信息学研究所大语言模型研发中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究复现比较多种提升语言模型提示鲁棒性的训练策略,发现现有方法虽优于标准微调与上下文学习,但仍存在40%-57%的性能差距,且最新鲁棒性方法未优于每批次单模板训练的简单策略,揭示了参数梯度符号冲突等关键原因。
AI 中文摘要
尽管大型语言模型性能强劲,但它们对提示的表述仍高度敏感。现有研究通过优化数据构建或专用鲁棒性目标来解决这一问题。我们在受控条件下复现并比较这些策略,测量它们缓解模型提示敏感性的效果。研究发现,当前的鲁棒性微调方法优于标准微调与上下文学习,但最优到最差提示的性能差距仍高达40%-57%。此外,我们测试的最新鲁棒性增强方法——用于对比对齐的CoIN和用于一致性正则化的PPCL——通常无法优于最简单的数据构建策略:每批次训练一个模板。我们的诊断结果解释了这些现象:辅助目标会改变其惩罚的量,但无法泛化到该范围之外;同时,57%-64%参数的单模板梯度符号存在冲突,导致不同数据构建策略产生差异,混合表述的批次会迫使优化器协调相互竞争的更新,而非找到与提示无关的共享更新方向。
英文摘要
Despite their strong performance, large language models remain highly sensitive to prompt formulation. Prior work addresses this through refined data construction or through dedicated robustness objectives. We reproduce and compare these strategies under controlled conditions, and measure how effective they are in addressing models' prompt sensitivity. We find the current robustness fine-tuning methods improve over standard fine-tuning and in-context learning, but the best-to-worst prompt gap remains as high as 40-57% of performance. Moreover, the recent robustness-enhancing methods we test - CoIN for contrastive alignment and PPCL for consistency regularization - often fail to outperform the simplest data construction strategy: training on one template per batch. Our diagnostics explain these results. The auxiliary objectives move the quantity they penalize, but do not generalize beyond it. Additionally, data construction strategies differ due to the conflicting signs of per-template gradients on 57-64% of parameters. Thus, batches that mix formulations force the optimizer to reconcile competing updates instead of finding a shared, prompt-agnostic one.
Comments5 pages, 5 figures, 13 tables. Camera-ready version; Accepted to EMNLP 2026 Findings