发表机构
Cohere Labs; Laboratoire Hubert Curien; Bennett University; Western University(Cohere实验室; 于贝尔·居里实验室; 班尼特大学; 西安大略大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究指令微调对语言模型置信度与生成理由词汇多样性的影响,发现指令微调虽小幅降低校准度但显著改变置信度,且对理由多样性的影响存在差异,二者是指令微调的不同效应。
AI 中文摘要
经过指令微调的语言模型在各类生成任务中表现出色,但近期研究显示其存在言语化过度自信问题。在问答任务中,模型的言语化过度自信可能与生成的支持性理由的一致性相关。本文研究指令微调诱导模型置信度变化时,生成的答案理由的词汇多样性是否也发生相应变化。我们在问答基准上评估了三组匹配的基础模型与指令微调模型,发现尽管预测准确率变化有限、基于似然的校准度下降,但指令微调始终会改变答案置信度。其次,我们观察到指令微调对理由多样性的影响并非均匀:跨理由多样性始终下降,而表面词汇多样性在不同模型和基准中变化方向与幅度均不同。最后,我们发现控制答案选择和理由长度后,这些差异仍然存在,证实置信度与理由多样性捕捉到指令微调的不同影响。
英文摘要
Instruction-tuned language models achieve strong performance across a range of generation tasks but have recently been shown to exhibit verbalized overconfidence, which may manifest in less diverse supporting rationales for incorrect answers. However, whether such overconfidence is associated with rationale consistency remains an open question. In this paper, we study whether changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently increases answer confidence, despite limited changes in predictive accuracy, while degrading likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.