发表机构
RTX BBN Technologies(RTX BBN技术公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建arXiv作者语料库,发现导师与学生、学术兄弟姐妹的写作风格相似性显著高于随机同领域作者,且作者身份归因系统的错误多集中于这类亲属配对。
AI 中文摘要
随着作者身份归因系统越来越多地被用于检测代笔和AI生成的论文,其错误可能会对合法作者提出指控,这类系统假设每位作者的风格是独有的。然而,研究者在导师指导下学习,会继承导师的风格特征。我们从数学谱系项目图谱中构建了拥有≥2篇独著论文的arXiv作者语料库,共包含5803位作者和2501组真实的导师-学生配对。使用微调模型的嵌入向量,导师与学生之间的余弦距离比随机同领域作者近39.9%;两个开源编码器分别在12.6%和14.5%的水平上复现了该效应。学术兄弟姐妹(同一导师的两名学生,可能从未谋面)在8360组配对中,即使在不同机构学习,风格距离也近30.4%;仅共享机构和领域的配对则表现出可忽略的相似性。在针对同一语料库的闭集归因任务中,系统对真实作者的导师和学术兄弟姐妹的错误率是随机情况的11倍。
英文摘要
As authorship attribution systems are increasingly deployed to detect ghostwritten and AI-generated papers, their errors can support accusations against legitimate authors. These systems conflate stylistic similarity with individual identity. Researchers, however, study under advisors, and inherit their stylistic quirks. We build a corpus of arXiv authors with $\geq 2$ solo papers from the Mathematics Genealogy Project graph, giving $5{,}803$ total authors and $2{,}501$ ground-truth advisor-student pairings. Using embeddings from a fine-tuned model, advisors sit $39.9\%$ closer in cosine distance to their students than a random same-field author does. Using two open models, we reproduce the effect at $12.6\%$ and $14.5\%$. Academic siblings, two students of one advisor who may never have met, sit $30.4\%$ closer across $8{,}360$ pairs, even when they studied at different institutions. Pairs who share only institution and field show negligible similarity. Given a closed-set attribution task over the same corpus, the system's errors occur on the true author's advisor, student, academic sibling, or lab mate $11$ times more often than chance.