发表机构
Bundesdruckerei GmbH(德国联邦印刷公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对德国公共部门构建MÖVE评估框架,发现现有LLM基准不适用于本地场景,不同模型在能耗、提供商透明度、德国政党立场认知维度存在显著权衡,公共机构选LLM需结合治理要求而非仅性能排名。
AI 中文摘要
公共机构在选择适配自身特定场景的大语言模型(LLM)时面临持续挑战。然而现有基准实用性有限,因其主要反映英语语境和以美国为中心的场景,且往往仅评估任务性能。本文呈现MÖVE(面向德国公共部门的整体评估框架)的首批结果,考察三个极少被考量的治理维度:能耗、提供商透明度以及对德国政党立场的认知。研究结果显示存在显著权衡,无单一模型能在所有维度表现优异:估算能耗差异超60倍,且无法仅通过模型规模解释;信息披露程度随提供商系统变化;欧洲模型未展现更强的德国政党立场认知能力。因此公共机构的模型选择不能仅依赖性能排名,评估还应反映部署场景的治理要求。
英文摘要
Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of MÖVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.
CommentsAccepted as non-archival paper at Eval4SD (co-located with KONVENS 2026)