arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27603cs.CLcs.AI

当上下文误导时:大型语言模型中带有管辖权的上下文学习

When Context Misleads: In-context Learning with Jurisdiction in Large Language Models

Pei-lin Li, Qingle Liu, Junyang Feng, Siyu Li, Sunqi Fan, Xin-Sheng Chen, Shuojin Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对上下文学习忽视上下文权威导致易受误导的问题,提出管辖权上下文学习(J-ICL)框架,通过将上下文验证纳入训练目标,在四个模型上平均提升ICLEval 5.84个百分点,现实准确性9.20个百分点,并提高对欺骗性上下文的抵抗力。

中文摘要 AI 辅助

上下文学习(ICL)已成为现代大型语言模型部署的基石。然而,现有的ICL后训练方法存在一个关键盲点:它们擅长从演示中提取模式,却常常忽视上下文权威性,即判断上下文信息是否应支配最终答案的能力。为了对此能力进行基准测试,我们引入了FakeContextBench,其中包含七个领域的伪科学主张。我们对商业和开源模型的评估表明,仅靠大规模预训练不足以实现可靠的上下文权威判别。此外,普遍的ICL微调方法可能增加对误导性上下文的敏感性,相对于基础模型,现实准确性最多降低14.95个百分点。为了解决这一权衡,我们提出了管辖权上下文学习(J-ICL),一种将上下文验证纳入训练目标的后训练框架。在四个模型骨干上,J-ICL相对于相应基础模型平均提高了ICLEval 5.84个百分点,现实准确性提高了9.20个百分点。它还相对于MetaICL和符号调优平均将现实率提高了18.09个百分点。这些结果表明,ICL能力和对欺骗性上下文的抵抗力可以同时提高。基准可在https://github.com/peilin717/FakeContext-Bench获取。

英文摘要

In-Context Learning (ICL) has become a cornerstone of modern LLM deployment. However, existing ICL post-training methods have a critical blind spot: they excel at extracting patterns from demonstrations while often neglecting context authority, the ability to determine whether contextual information should govern the final answer. To benchmark this capability, we introduce FakeContextBench, which contains pseudoscientific claims across seven domains. Our evaluation of commercial and open-source models shows that large-scale pre-training alone is insufficient for reliable context-authority discrimination. Moreover, prevalent ICL fine-tuning methods can increase susceptibility to misleading context, reducing reality accuracy by up to 14.95 percentage points relative to the base model. To address this trade-off, we propose Jurisdiction In-Context Learning (J-ICL), a post-training framework that incorporates context validation into the training objective. Across four model backbones, J-ICL improves ICLEval by an average of 5.84 percentage points and reality accuracy by 9.20 points over the corresponding base models. It also raises the Reality Rate by an average of 18.09 points relative to MetaICL and Symbol Tuning. These results demonstrate that ICL capability and resistance to deceptive context can be improved together. The benchmark is available at https://github.com/peilin717/FakeContext-Bench.

发表机构

  • Tsinghua University(清华大学)
  • Huazhong University of Science and Technology(华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑