arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16627cs.CLcs.AI

解释何时能帮助上下文学习?自然语言解释类型与忠实度的比较研究

When Do Explanations Help In-Context Learning? A Comparative Study of Natural Language Explanation Types and Faithfulness

  • LMU Munich(慕尼黑大学)
  • Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
  • Imperial College London(伦敦帝国学院)
  • Technical University of Munich(慕尼黑工业大学)
  • MaiNLP lab, CIS, LMU Munich(慕尼黑大学CIS研究所MaiNLP实验室)

机构由 AI 辅助整理,请以论文原文为准。

Mahdi Dhaini, Adam Dejl, Juraj Vladika, Volkan Özer, Barbara Plank, Gjergji Kasneci

AI总结:

该研究通过在6个基准和4个指令调优模型上的比较评估,探究不同来源与选择方式的自然语言解释对上下文学习下游性能的影响,为实际提示流程中解释的选择和报告提供了见解。

AI中文摘要:

自然语言解释(NLEs)正日益被用作输入,例如作为影响上下文学习(ICL)中模型行为的少样本基本原理。然而,在解释增强提示中,不同类型的NLE对下游模型性能的影响如何比较仍不清楚。因此,我们在6个基准和4个指令调优模型上进行了比较评估,研究NLE来源(可用时为人类编写的、自生成的解释、由外部大语言模型生成的解释)和NLE选择(随机选择与基于忠实度的过滤)在ICL设置中使用时对NLE下游效用的影响。我们的广泛评估表明,在分类式基准上,将NLE添加到少样本提示中通常比无解释的少样本提示能提高准确率;在NLE来源中,外部生成的LLM-NLE通常提供较强的下游效用,在两者都可用时仍可与人类基本原理相媲美,而自生成的NLE(self-NLE)对选择策略更敏感。在数学推理任务上,效果更依赖于模型和来源。我们进一步表明,基于忠实度选择自生成的NLE总体上能产生小的平均增益,但根据指标、任务和模型的不同,它可能会提高或降低性能。不同的忠实度指标可能存在显著分歧,这会影响选择哪些自生成的NLE示例及其下游预测效用。通过随机交换和分布外基本原理进行的稳健性测试表明存在部分稳健性,这表明语义对齐有助于性能提升。总体而言,我们的结果为在实际提示流程中选择和报告影响模型行为的解释提供了见解。

英文摘要:

Natural language explanations (NLEs) are increasingly used as inputs, for example, as few-shot rationales that influence model behavior in in-context learning (ICL). However, it remains unclear how different types of NLEs compare in their effects on downstream model performance in explanation-augmented prompting. Therefore, we provide a comparative evaluation across six benchmarks and four instruction-tuned models, studying how NLE source (human-written when available, self-generated explanations, generated by an external LLM) and NLE selection (random vs faithfulness-based filtering) affect downstream utility of NLEs when used in ICL settings. Our extensive evaluation shows that, on classification-style benchmarks, adding NLEs to few-shot prompts often improves accuracy over few-shot prompting without explanations; among NLE sources, externally generated LLM-NLEs often provide strong downstream utility and remain competitive with human rationales where both are available, whereas self-NLEs are more sensitive to the selection strategy. On math reasoning, the effects are more model- and source-dependent. We further show that faithfulness-based selection of self-NLEs yields small average gains overall, but can improve or reduce performance depending on the metric, task, and model. Different faithfulness metrics can disagree substantially, affecting which self-NLE examples are selected and their downstream predictive utility. Robustness tests with randomly swapped and out-of-distribution rationales indicate partial robustness, suggesting that semantic alignment contributes to performance gains. Overall, our results provide insights for selecting and reporting explanations that influence model behavior in practical prompting pipelines.

↑