发表机构
School of Computing and Information Systems, The University of Melbourne; University of Calabria(墨尔本大学计算与信息系统学院; 卡拉布里亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MultiGhostBench是含928本多语言长文本的LLM生成文本基准,支持多类分布偏移评估,可用于开发鲁棒的LLM作者身份归因方法。
AI 中文摘要
尽管现有的大语言模型(LLM)作者身份归因(AA)研究已取得一定进展,但可用基准仍存在局限,通常聚焦于英语、受控设置或相对过时的模型,少数多语言研究仅考虑较短文本。我们推出MultiGhostBench,这一多语言基准包含5种近期LLM生成的、覆盖6种语言和3种书写体系的928本书籍,每本书平均长度约为5.9万字。该基准支持在领域、作者及语言偏移下开展评估。对代表性AA方法的评估显示,没有任何一种方法能在所有设置下始终表现最佳,且性能通常会在分布偏移下下降。基于Transformer的检测器可跨语言保留与生成器相关的信息,不过迁移效果因语言对而异;而基于统计和指纹的检测器则更依赖语言。我们预计MultiGhostBench将成为开发和评估鲁棒AA方法的宝贵资源。数据集和代码可在指定网址获取。
英文摘要
While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts. We introduce MultiGhostBench, a multilingual benchmark comprising 928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 59K words per book. The benchmark supports evaluation under domain, author, and language shifts. Evaluation of representative AA methods shows that no single method consistently performs best across settings, and performance generally degrades under distribution shifts. Transformer-based detectors can retain generator-related information across languages, although transfer effectiveness varies by language pair, whereas statistical and fingerprint-based detectors are more language-dependent. We envision MultiGhostBench as a valuable resource for the development and evaluation of robust AA methods. The dataset and code can be found at https://github.com/GrecoMT/MultiGhostBench.