大语言模型作为数据标注器的有效性:关于权威的AMALIA研究
Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
浏览论文内容
中文总结 AI 辅助
研究葡萄牙AMALIA大语言模型作为数据标注器标注权威道德基础的有效性,通过恢复差距测试其能否遵循理论,发现校准的英语工具不能转移到该模型,其依赖表面关联,虽能筛选预编码但不能独立衡量,强调基准测试应考察一致性证据路径。
中文摘要 AI 辅助
国家语言模型为语言社区提供了衡量公民言论和价值观的工具。葡萄牙的AMALIA是一个由公共资金支持的、拥有90亿参数的欧洲葡萄牙语模型,仅在一致性方面就颇具竞争力。然而,一致性是可靠性而非有效性。对于必须从表面特征推断而非直接读取的理论结构,问题在于模型是遵循结构的理论还是通过相关捷径得出正确代码。我们用恢复差距来测试这一点。我们询问经过校准的英语工具是否能转移到AMALIA - 9B和欧洲葡萄牙语中。对于一个结构和一个语料库,结果是否定的。分解仅恢复了AMALIA整体性能的约一半,错误分析表明其依赖表面关联。一个开放的多语言大语言模型在相同指令下缩小了差距,这表明问题不在于语料库。AMALIA仍可大规模筛选和预编码,但还不能很好地衡量该结构以独立使用。该研究虽非对国家模型的定论,但认为主权大语言模型基准测试不仅应测试与人类编码者的一致性,还应测试达成该一致性的证据路径。
英文摘要
National language models are becoming publicly funded epistemic infrastructure. Public ownership, linguistic specialization, and open weights create a presumption of trustworthiness. Such an instrument, built by and for a language community, looks like the natural choice for measuring what that community says and values. Whether such a model validly measures anything is untested at release. The evaluation of LLMs as measurement instruments is typically task-specific and stops at agreement with human coders. Agreement cannot distinguish an LLM instrument that measures a construct from one that reaches matching codes through surface correlates. We audit the presumption on a favourable case: AMALIA, Portugal's publicly funded 9B model, coding the moral foundation of authority in European Portuguese. The \textit{recovery gap} operationalizes the audit: decompose the codebook into its theory-defined clauses, recombine them through the theory's explicit rule, and measure how much of the original prompt's performance the stated theory reproduces. In a pre-registered, out-of-sample study on a transcreated (English to European Portuguese) corpus, AMALIA agrees with trained coders within six points of open models eight to thirteen times its size. Yet, the recovery gap shows that only about half of coding performance on authority can be attributed to the theory. A larger multilingual LLM closes the recovery gap on the same corpus, suggesting the shortfall lies in the annotator model, not the corpus or its translation. Sovereignty earns operational and performance trust; epistemic trust requires calibration -- and the audit method is inexpensive, and portable across models, languages and tasks.
发表机构
- CICANT, Universidade Lusófona(坎特交互计算与技术研究中心,卢索福纳大学)
机构由 AI 辅助整理,请以论文原文为准。