发表机构
The National Center for AI in Society (CAISA), University of Copenhagen(哥本哈根大学社会人工智能国家中心(CAISA))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM政治立场检测中模板化提示的混淆问题,扩展IssueBench框架,提出用LLM生成的完全合成提示,验证其在立场测量中更具生态效度,可减少估计偏差。
AI 中文摘要
长期以来,大型语言模型(LLM)的政治立场检测一直被封闭式多项选择政治调查问题主导,这些问题最初是为人类设计的,因此缺乏现实世界人机交互的真实性和细微差别,同时还容易出现“放水(sandbagging)”现象。近期的IssueBench框架通过基于真实聊天记录的模板化提示,大幅缓解了这些局限性。鉴于生成式AI(GenAI)助手在非工作相关场景中的使用日益增多,我们将IssueBench的应用范围从写作辅助扩展到另外两项任务:信息检索和观点分享。我们认为,模板化提示仍然缺乏真实提示的细微差别,尤其是在开放式任务中,且仍可被识别为评估人工制品。我们提出使用完全合成的(由LLM生成的)提示,这些提示是在详细指令下以真实提示为种子生成的。我们在一项小规模研究中评估了真实提示、模板化提示和LLM生成提示的生态效度,该研究涵盖3个极具争议的政策问题和3个近期地缘政治冲突。人类和LLM标注者认为,LLM生成的提示的真实性不低于真实提示,且明显比模板化提示更真实,还发现它们能更清晰地传递预期意图和立场;LLM将模板化提示与另外两类提示区分的清晰度远高于人类。在案例研究中,对于同一模型,模板化提示和LLM生成提示产生的立场估计存在系统性差异,在中立框架下最为明显,其中模板化提示夸大了模型在模板中插入的主题与立场文本(填充内容)所编码方向上的倾向。
英文摘要
Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbagging. The recent IssueBench framework substantially mitigates these limitations with templated prompts anchored in real-world chat logs. Given the rise in non-work-related use of GenAI assistants, we extend IssueBench beyond writing assistance to include two additional tasks, information seeking and opinion sharing. We argue that templated prompts still lack the nuance of real ones, especially for open-ended tasks, and remain recognisable as evaluation artefacts. We propose the use of fully synthetic (LLM-generated) prompts, produced under detailed instructions with real prompts as seeds. We assess the ecological validity of real, templated, and LLM-generated prompts in a small-scale study covering 3 highly contested policy issues and 3 recent geopolitical conflicts. Human and LLM annotators rank LLM-generated prompts as no less realistic than real ones and clearly more realistic than templated ones, and find that they carry their intended intent and stance more clearly; the LLMs separate templated prompts from the other two far more sharply than the humans do. In a case study, templated and LLM-generated prompts yield systematically different stance estimates for the same model, most visibly under neutral framings, where templated prompts overstate the model's leaning in the direction encoded by the topic-and-stance text (filler) slotted into their templates.