arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从价值到基准:评估用于荷兰政府场景的大型语言模型

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

Laurens Samson, Iva Gornishka, Gossa Lô, Yuki M. Asano, Sennay Ghebreab

arXiv 2608.09925首次发表:更新:

发表机构

City of Amsterdam; University of Amsterdam; University of Technology Nuremberg(阿姆斯特丹市政府; 阿姆斯特丹大学; 纽伦堡工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出面向荷兰政府场景的“Grip on LLMs”评估框架,确定六大评估维度,发现模型在质量、成本、能耗等维度存在权衡,发布模型概览供政府LLM选择使用。

AI 中文摘要

大型语言模型正越来越多地被部署到政府场景中,但现有的评估框架很少能同时反映公共行政的价值以及非英语语境的语言要求。我们与荷兰某大型市政机构的领域专家合作,推出了“Grip on LLMs”框架,这是一套用于荷兰政府场景的系统性评估套件。通过顾问委员会流程、用户研究以及对公务员聊天机器人用户的调查,我们确定了六个评估维度:事实性、诚实性、社会偏见、能耗、成本和训练数据透明度,并将其转化为覆盖30多个多语言模型和荷兰特定模型的基准套件。我们的结果显示,没有任何单一模型能在所有维度上表现出色,权衡取舍不可避免:更高的质量总是伴随着更大的环境影响和财务成本,而偏见则在很大程度上与这两者无关。我们还发现,事实性(模型是否正确回答问题)和诚实性(模型是否承认自身的未知内容)受不同属性支配,高事实性并不意味着高诚实性。为了让这些发现对非技术受众具有可操作性,我们发布了一个公开可访问、用户友好的模型概览,面向参与政府LLM选择的所有利益相关方,从工程师到政策制定者。

英文摘要

Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions (factuality, honesty, social bias, energy consumption, cost, and training data transparency) and operationalise them into a benchmark suite covering more than 30 multilingual and Dutch-specific models. Our results reveal that no single model excels across all dimensions, and that trade-offs are unavoidable: higher quality consistently comes at greater environmental impact and financial cost, while bias remains largely independent of both. We further find that factuality (whether a model answers correctly) and honesty (whether a model acknowledges what it does not know) are governed by distinct properties, with high factuality not implying high honesty. To make these findings actionable for non-technical audiences, we release a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in governmental LLM selection, from engineers to policymakers.

CommentsAccepted at AIES 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑