可复现性并非构念效度:LLM 对制度情境化沟通的测量
Reproducibility is not construct validity: LLM measurement of institutionally situated communication
- Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院)
- Inria(法国国家信息与自动化研究所)
- Sciences Po(巴黎政治学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究利用欧盟AI法案咨询数据,发现LLM注释虽高度可复现但与构念效度脱节,且分歧因利益相关方而异,凸显了区分可复现性与构念效度的必要性。
AI中文摘要:
高注释可复现性并不必然意味着 LLM 推断出的测量指标捕捉到了其旨在测量的构念。我们利用欧盟委员会《人工智能法案》咨询中的数据集检验了这一区别,将结构化调查回复与来自同一利益相关方的自由文本咨询提交相链接。LLM 对咨询提交的注释高度可复现(组内相关系数 > 0.99),但与其旨在近似的名义构念的调查报告测量值之间仅显示出有限的收敛性。调查推断与 LLM 推断的基于文本的测量值之间的分歧在不同利益相关方群体中系统性变化:商业协会在基于文本的咨询中对 AI 风险表达了比调查回复更大的担忧(g = +1.0),而公共当局和几个非商业群体则显示出较小或负向的分歧。分数之间的分歧表明欧洲国家之间存在正向空间自相关(Moran's I = 0.347,p = 0.036),这表明来自邻国的利益相关方倾向于对 AI 安全问题采取更相似的基于文本的立场。尽管存在分歧,但在所有分歧水平上,调查报告的担忧仍与对可解释性的支持密切相关。这些结果表明,LLM 注释的可复现性可以与较差的构念对应性共存,并促使验证程序区分可复现性、构念效度和沟通情境变化,当 LLM 被用作测量工具时。
英文摘要:
High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.