arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EuroExec:前沿语言模型在欧洲行政决策任务上未能达到专家判断水平

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks

Pau Arnal, Khaled Denfir, Danylo Smahliuk, Amrut Avhad, Marcus A. Castro

arXiv 2608.04549首次发表:更新:

发表机构

Sovrano AI(索夫拉诺人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建含413项欧洲行政任务的EuroExec基准,评估6款前沿LLMs,发现其解决率仅56.9%,远低于专家水平,凸显前沿LLMs在这类任务上的不足及人工评估的必要性。

AI 中文摘要

前沿大语言模型(LLMs)正越来越多地被用于解决开放式复杂问题,这类问题与它们通常被评估的问题性质不同。我们投入了超过4000小时的人类专家时间,评估了6款前沿LLMs在这类问题中的一个子集上的表现:EuroExec,这是我们提出的、基于人类专家的基准,由47名经过审核的领域专家编写的413项开放式长格式欧洲行政任务构成,每项问题均来自实际案例经验。所有回答均通过多属性评分标准、针对具体条目的要求清单以及偏好排名进行人工评估,提取出汇总指标“解决率”。性能最强的模型仅能解决56.9%的任务,而盲审的专家编写参考答案的解决率接近最高水平,且在74%的直接排名中被优先于所有模型回答,这表明前沿生成系统的表现远低于它们已被应用的专业工作标准。我们认为,得出此类结论的最佳方式是聘请人类评估者,通过严格的统计分析仔细检查其一致性,同时观察到在评估这类具有主观真实情况的现实世界开放式问题时,自动测量方法也存在不足。

英文摘要

Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated on. We dedicate more than 4,000 human expert hours to evaluate a selection of six frontier LLMs on a member of this class of problems: EuroExec, our introduced human expert-based benchmark composed of 413 open-ended long-form European executive tasks authored by 47 vetted domain experts, each question drawn from experience in a real case. Every response is manually evaluated through a multi-attribute rubric, an item-specific checklist of requirements, and a preference rank ordering, extracting an aggregate metric "Solve Rate". The strongest model solves only 56.9% of tasks, while expert-written reference answers judged blindly are solved at near-ceiling levels and are preferred over every model response in 74% of direct rankings, placing frontier generative systems well below the professional standard of work they are already used for. We see that the best way to extract this kind of conclusion is by employing human evaluators, carefully checking their consistency through rigorous statistical analysis, and observe that automatic measurements also fall short when evaluating on this case of real-world open-ended problems with a subjective ground truth.

Comments17 pages, 9 figures, 12 tables, submitted to EACL 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑