arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过还是失败?评估大语言模型在两个希腊考试基准上的表现

Pass or Fail? Evaluating LLMs on Two Greek Examination Benchmarks

Panagiota Kyriazi, Eleni Kasoura, Prokopis Prokopidis

arXiv 2609.34800首次发表:更新:

发表机构

Institute for Language and Speech Processing / Athena RC(语言与语音处理研究所 / 雅典研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出 Prot-Ex 和 Pan-Ex 两个希腊考试基准,评估多种 LLMs,发现本地化 KriKri-8B 表现优异,并揭示少样本提示在结构化任务中的上下文过载悖论。

AI 中文摘要

大语言模型(LLMs)的快速发展要求对其语言和分析能力及局限性进行彻底评估,特别是对于像希腊语这样基准覆盖有限的语言。为了解决该领域全面基准稀缺的问题,我们引入了 Prot-Ex 和 Pan-Ex 两个基准,它们包含来自希腊示范学校和实验学校入学考试以及泛希腊考试(希腊国家大学入学考试)的题目。这些基准被用于评估纯文本 LLMs 的性能,包括希腊语适配的 KriKri-8B-Instruct、Llama-3.1-8B、Gemma-4-26B 和 Qwen-3-32B,涵盖多种学科(现代希腊语、数学、物理等)和任务形式(封闭式、结构化和开放式),以及文本化的视觉上下文(即图像描述)。我们的研究结果表明,本地化的 KriKri-8B 显著优于其基础模型,在语言要求较高的人文学科任务中成功与更大的 LLMs 抗衡。通过利用 LLM-as-a-Judge 方法,我们揭示了传统词汇指标在评估复杂推理方面的不足。关键的是,我们发现了一个少样本提示悖论:虽然合成示例提高了封闭式问题的准确性,但在结构化任务中严重过载了 8B 模型的上下文窗口,导致显著的性能下降。最终,这项研究表明,针对性的语言适配可以弥补较小参数数量在专业领域的不足,尽管较小模型对提示冗长性较为脆弱。

英文摘要

The rapid advancement of Large Language Models (LLMs) imposes a thorough evaluation of their linguistic and analytical capabilities as well as constraints, particularly for a language with limited benchmark coverage such as Greek. To address the limited availability of comprehensive benchmarks in this domain, we introduce Prot-Ex and Pan-Ex, two benchmarks consisting of questions from entrance exams for Greek Model and Experimental schools as well as the Panhellenic exams (the Greek national university entrance examinations). These benchmarks are employed to assess the performance of text-only LLMs-including the Greek-adapted KriKri-8B-Instruct, Llama-3.1-8B, Gemma-4-26B, and Qwen-3-32B-across diverse academic disciplines (Modern Greek, Mathematics, Physics, etc.) and task formats (closed, structured, and open-ended), including textualized visual context (i.e., image descriptions). Our findings indicate the localized KriKri-8B significantly outperforms its base model, successfully rivalling much larger LLMs in linguistically demanding humanities tasks. By leveraging an LLM-as-a-Judge methodology, we expose the inadequacy of traditional lexical metrics for evaluating complex reasoning. Crucially, we uncover a few-shot prompting paradox: while synthetic examples improve accuracy in closed-ended questions, they severely overload the context window of 8B models in structured tasks, causing significant performance degradation. Ultimately, this study suggests targeted linguistic adaptation offsets lower parameter counts in specialized domains, despite the fragility of smaller models to prompt verbosity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑