arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11598cs.IR

NTCIR-19大语言模型自动评估2(AEOLLM-2)任务概述

Overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) Task

Junjie Chen, Yuxi Dong, Haitao Li, Yiqun Liu, Qingyao Ai

首次发表
浏览论文内容

中文总结 AI 辅助

本文概述NTCIR-19的AEOLLM-2任务,该任务在AEOLLM基础上新增深度研究评估子任务,共收到10个团队的91次运行结果,介绍了任务背景、数据集构建等内容。

中文摘要 AI 辅助

本文概述了NTCIR-19大语言模型自动评估2(AEOLLM-2)任务。基于NTCIR-18核心任务AEOLLM的成功,我们为NTCIR-19推出AEOLLM-2,以进一步研究大语言模型(LLM)的自动评估方法,尤其聚焦于长文本生成场景。在AEOLLM-2中,我们新增了深度研究评估子任务,专注于自动评估LLM生成的长文本深度研究报告。参与者开发自动评估这些报告质量的方法,各方法的性能通过将其评分与人工标注的基准真值标签对比来衡量。本年度我们共收到10个团队提交的91次运行结果。本文介绍了该任务的背景、数据集构建、评估指标、参与者方法及最终评估结果。

英文摘要

In this paper, we provide an overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) task. Building on the success of the NTCIR-18 core task AEOLLM, we proposed AEOLLM-2 for NTCIR-19 to further investigate automatic evaluation methods for Large Language Models (LLMs), particularly in long-form text generation scenarios. In AEOLLM-2, we introduced a new subtask, Deep Research Evaluation, which focuses on the automatic evaluation of long-form deep research reports generated by LLMs. Participants developed evaluation methods to automatically assess the quality of these reports, and the performance of each method was measured by comparing its scores against human-annotated ground-truth labels. This year, we received 91 runs from 10 teams in total. This paper describes the background of the task, the dataset construction, the evaluation measures, the participants' methods, and the final evaluation results.

发表机构

  • DCST, Tsinghua University , Quan Cheng Laboratory(清华大学计算机科学与技术系,量子科学与工程研究院)
  • University of Science and Technology Beijing(北京科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑