学习失败的诱因:用于对抗性数据策展的失败模式上下文多臂老虎机
Learning What to Fail On: Failure-Mode Contextual Bandits for Adversarial Data Curation
浏览论文内容
中文总结 AI 辅助
该研究提出失败模式上下文多臂老虎机框架,通过检索增强等技术优化对抗性数据策展,在多项自然语言理解基准上提升模型准确率,还可迁移至事实核查任务,实现无需额外人工标注的鲁棒性提升。
中文摘要 AI 辅助
我们提出了一种感知失败的对抗性检索增强框架,以提升自然语言理解的鲁棒性。与采用固定奖励阈值选择合成示例不同,我们的方法将对抗性数据策展建模为失败模式上下文多臂老虎机问题。候选示例通过检索增强提示生成,经当前目标模型过滤,由大型语言模型(LLM)评判集成自动验证,并聚类为重复出现的失败模式。随后,随机策略选择需采样以进行再训练的失败模式,并使用基于验证的奖励进行更新,该奖励在鲁棒性提升、遗忘效应与数据成本之间取得平衡。这使得数据策展器本身成为学习智能体,能够在多轮训练中自适应选择最有用的模型失败模式。在标准基准测试中,我们的方法将RoBERTa-base在SNLI上的准确率从88.48%提升至92.60%,在ANLI上从75.04%提升至80.95%,在MultiNLI上从54.67%提升至71.99%,且持续优于现有对抗性增强方法。我们进一步验证了其在FEVER事实核查任务中的迁移能力,使用RoBERTa-large最高可实现79.86%的FEVER分数与82.45%的准确率。最后,我们提供了理论解释,表明在给定假设下,失败模式采样可减少捷径对齐的梯度贡献,同时诱导有界分布漂移。通过结合检索、自动验证、上下文多臂老虎机的失败选择与受控对抗性再训练,我们的框架无需额外人工标注即可实现可扩展的鲁棒性提升。
英文摘要
We introduce a failure-aware adversarial retrieval-augmented framework for improving robustness in natural language understanding. Rather than selecting synthetic examples with a fixed reward threshold, our method formulates adversarial data curation as a failure-mode contextual bandit problem. Candidate examples are generated with retrieval-augmented prompting, filtered by the current target model, automatically validated by an LLM judge ensemble, and clustered into recurring failure modes. A stochastic policy then selects which failure modes to sample for retraining, and is updated using validation-based reward that balances robustness gains, forgetting, and data cost. This makes the data curator itself the learning agent, enabling adaptive selection of the most useful model failures across training rounds. On standard benchmarks, our approach improves RoBERTa-base accuracy from 88.48% to 92.60% on SNLI, from 75.04% to 80.95% on ANLI, and from 54.67% to 71.99% on MultiNLI, while consistently outperforming prior adversarial augmentation methods. We further demonstrate transfer to FEVER fact verification, achieving up to 79.86\% FEVER score and 82.45\% accuracy with RoBERTa-large. Finally, we provide a theoretical interpretation showing that, under stated assumptions, failure-mode sampling can reduce shortcut-aligned gradient contributions while inducing bounded distributional drift. By combining retrieval, automated validation, contextual-bandit failure selection, and controlled adversarial retraining, our framework enables scalable robustness improvement without additional human annotation.
发表机构
- Ben Gurion University of The Negev(本古里安大学内盖夫分校)
机构由 AI 辅助整理,请以论文原文为准。