arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MORFES:现代希腊语屈折生成能力基准测试集

MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

Ioakeim Perros, Cleopatra Papadopoulou, Ayoub Kirouane, Christos Petrocheilos

arXiv 2607.28274首次发表:更新:

发表机构

Sophea AI(Sophea AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现代希腊语屈折生成能力缺乏专用基准的问题,构建了含500个条目的MORFES基准,评估多款开源模型,其中自研的Sophea-Genesis-1在屈折形态任务中表现领先。

AI 中文摘要

现代希腊语是屈折变化丰富的语言,但针对该语言构建的语言模型主要基于事实知识进行评估,目前尚无专门针对其屈折生成能力的基准测试集。我们推出了MORFES(形态开放类识别与生成评估套件,Morphological Open-class Recognition-and-Formation Evaluation Suite),这一基准测试集包含500项经专家验证的条目,用于测试希腊语屈折形式的识别与生成能力,优先选用低频词元,确保正确答案反映的是规则而非记忆的形式。我们将其公开于该https URL。我们在MORFES上评估了一系列开源语言模型,将其置于从LLaMA到Qwen3、DeepSeek-R1、Magistral、Kimi K2快速扩展的开源权重生态系统中,该系统的多语言覆盖范围不断扩大,但对屈折变化丰富语言的语法能力测量仍显不足。其中,我们开发并以开源权重发布的模型Sophea-Genesis-1在屈折形态任务中表现领先,同时在通用能力上与同规模模型相当。

英文摘要

Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicated to their inflectional competence. We introduce MORFES (Morphological Open-class Recognition-and-Formation Evaluation Suite), a benchmark of 500 expert-verified items that tests the recognition and production of Greek inflected forms, favoring lower-frequency lemmas so that a correct answer reflects the rule rather than a memorized form. We make it publicly available at https://huggingface.co/datasets/KIEFERSA/MORFES. We evaluate a range of open language models on MORFES, situating them within the rapidly scaling open-weight ecosystem from LLaMA to Qwen3, DeepSeek-R1, Magistral, and Kimi K2, where multilingual coverage grows but grammatical competence in morphologically rich languages remains under-measured. Among them, Sophea-Genesis-1, a model we developed and release as open weights at https://huggingface.co/KIEFERSA/Sophea-Genesis-1, leads on inflectional morphology while matching similarly sized models in general capability.

Comments12 pages, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑