发表机构
Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AuthBench是一个大规模多语言作者身份表征基准,涵盖多种语言、体裁和长度,通过统一零样本协议评估47个神经模型,发现作者身份表征仍远未解决,并提供诊断资源以研究其成败原因。
AI 中文摘要
作者身份信号在写作风格承载身份的场合中至关重要:数字取证、剽窃分析、账号关联、虚假信息调查以及机器生成文本检测。然而,当前的作者身份基准仍然零散,通常仅覆盖狭窄的语言集合、单一体裁或有限的文档长度范围,这使得评估现代表征是否真正具备泛化能力变得困难。我们提出了AuthBench,一个用于作者身份表征的大规模多语言基准,旨在使这一评估变得广泛、标准化且贴近现实。AuthBench包含由153,825名作者撰写的428,150篇文档,涵盖十种广泛使用的语言、9种主要体裁、66种细粒度体裁以及四个文档长度区间。它支持两个互补的任务:作者归属,形式化为同作者检索;以及作者验证,形式化为同作者二元决策。我们在统一的零样本协议下对47个神经模型和三个非神经基线进行了基准测试。结果表明,作者身份表征仍远未解决:最佳检索模型仅达到0.258的Success@5,而最佳验证模型实现了0.076的EER和0.968的ROC-AUC。排行榜还揭示了有意义的任务分化,不同的模型家族分别在检索和验证中领先,并且在语言、体裁和长度上存在显著的性能差异。这些发现使AuthBench不仅作为一个新基准,而且作为研究作者身份表征何时以及为何成功或失败的诊断资源。我们在https URL和https URL发布了AuthBench、其评估工具包和基准数据。
英文摘要
Authorship signals matter in settings where writing style carries identity: digital forensics, plagiarism analysis, account linking, misinformation investigation, and machine-generated text detection. Yet current authorship benchmarks remain fragmented, usually covering only a narrow language set, a single genre, or a limited document-length regime, which makes it difficult to assess whether modern representations truly generalize. We introduce AuthBench, a large-scale multilingual benchmark for authorship representation that is designed to make this evaluation broad, standardized, and realistic. AuthBench contains 428,150 documents written by 153,825 individuals across ten widely used languages, 9 primary genres, 66 fine-grained genres, and four document-length buckets. It supports two complementary tasks: authorship attribution, formulated as same-author retrieval and authorship verification, formulated as same-author binary decision. We benchmark 47 neural models and three non-neural baselines under a unified zero-shot protocol. Results show that authorship representation remains far from solved: the best retrieval model reaches only 0.258 Success@5, while the best verification model achieves 0.076 EER and 0.968 ROC-AUC. The leaderboard also reveals a meaningful task split, with different model families leading retrieval and verification, and large performance differences across languages, genres, and lengths. These findings position AuthBench not only as a new benchmark, but as a diagnostic resource for studying when and why authorship representations succeed or fail. We release AuthBench, its evaluation toolkit, and benchmark data at https://github.com/mao-code/AuthBench and https://huggingface.co/datasets/MaoXun/AuthBench.