发表机构
Brock University; Emory University; University of Southern California; University College London; Massachusetts Institute of Technology(布鲁克大学; 埃默里大学; 南加州大学; 伦敦大学学院; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究通过控制实验,让参与者使用不同宣传话术下的六种大语言模型完成任务,发现宣传话术影响用户评价和交互行为,任务表现未变,用户评价更多反映期望管理。
AI 中文摘要
想象两个用户与同一大语言模型交互,被告知不同模型信息后评价差异大。在控制研究中,162名参与者使用六种大语言模型完成三项协作任务,预交互框架改变用户意见和行为,任务表现不变。用户评价更多取决于期望而非模型性能,使用后的印象变化也取决于期望和信心。
英文摘要
Imagine two users interact with the same LLM. One has been told it is the cutting-edge flagship model; the other, an older, weaker model. They walk away with markedly different ratings of its usefulness and intelligence, yet they used the same model. In a controlled study, 162 participants each used one of six LLMs from two families across three collaborative tasks, after first viewing a landing page that matched, overstated, or understated their model's true capability. This pre-interaction framing shifted user opinions and interaction behavior while task performance did not. Oversold users rated the model more favorably and used more directive prompting, while Undersold users wrote longer, more collaborative prompts. The quality of what users and the model produced together depended only on the model's true capability, not on what users were told. Participants' change in model impressions after use, measured across two impression measures, was not predicted by task performance ($β= -0.01$ and $0.11$, both n.s.), but by whether the model met users' expectations ($β= 0.47$ and $0.50$, both $p < .001$) and how confident they felt working with it ($β= 0.47$ and $0.36$, both $p < .001$). After interaction, users are still rating the pitch, not the product: user-elicited LLM evaluations, including the preference data driving public leaderboards, measure expectation management at least as much as the model itself.
CommentsAccepted to EMNLP 2026 Main Conference