arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提示鲁棒性取决于任务:在大语言模型评估中比较客观和信念风格问题

Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

Sadia Kamal, Arefa Patwary, Anthony Marchiafava, Sagnik Ray Choudhury, Atriya Sen

arXiv 2607.05554首次发表:更新:

发表机构

Oklahoma State University; University of North Texas(俄克拉荷马州立大学; 北德克萨斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型评估中提示鲁棒性,比较客观与主观问题,在多个数据集上评估四个模型家族,通过多种提示更改及特定方程分析,发现提示鲁棒性取决于问题类型、提示变化和模型。

AI 中文摘要

对大语言模型的调查式评估通常将提示响应视为模型价值观或信念的一种衡量标准。当将响应视为政治价值观、社会态度或信念的证据时,这一假设尤其脆弱。我们探讨了具有固定答案的客观问题和询问意见或价值观的主观问题之间的提示鲁棒性是否存在差异。我们在三个客观数据集(MMLU、ARC和CulturalBench)和三个主观数据集(政治指南针测试、价值基准和世界价值观调查)上评估了四个指令微调模型家族。对于每个问题/陈述,我们应用多种类型的提示更改,如措辞、框架和格式的变化,并衡量模型在不同变体中是否给出相同答案。使用二项式广义估计方程,我们发现模型、数据集、提示类别及其交互作用具有显著影响。数据集类型效应也很显著,数据集类型和提示类别之间的交互作用很大。这些结果表明,提示鲁棒性取决于问题类型、提示更改和模型。

英文摘要

Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs. This assumption is particularly fragile when responses are read as evidence of political values, social attitudes, or beliefs. We ask whether prompt robustness differs between objective questions with fixed answers and subjective questions that ask for opinions or values. We evaluate four instruction-tuned model families on three objective datasets (MMLU, ARC, and CulturalBench) and three subjective datasets (Political Compass Test, ValueBench, and World Values Survey). For each question/statement, we apply multiple types of prompt changes, such as variations in wording, framing, and format, and measure whether the model gives the same answer across variants. Using a binomial generalized estimating equation, we find significant effects of model, dataset, prompt category, and their interactions. The dataset type effect is also significant, and the interaction between dataset type and prompt category is large. These results show that prompt robustness depends on the question type, the prompt change, and the model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑