超越表征相似性:面向生成式剽窃检测与候选源重排序的源条件描述长度增益
Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking
浏览论文内容
中文总结 AI 辅助
本研究针对生成式剽窃检测难题,提出源条件描述长度增益(SCDG)框架,在PAN基准等测试中优于现有基线,可有效检测大量改写后的源复用。
中文摘要 AI 辅助
大型语言模型(LLMs)对学术诚信和同行评审构成挑战,但生成式剽窃检测仍是未被充分探索且基本未解决的难题。现有针对LLM生成文本检测的工作聚焦于AI参与情况,而这可能是可允许的,并非针对源复用;基于相似性的方法在经过大量改写和多源合成后效果不佳。受概率预测的描述长度观点启发,相关辅助信息可缩短目标序列的编码长度,我们提出源条件描述长度增益(SCDG),这是一种定向的无训练框架,将冻结语言模型在可疑文档P包含候选源S和不包含候选源S时的描述长度进行对比。该对比产生的词级对数似然增益可衡量S提供的增量预测证据。我们在CLEF的PAN生成式剽窃基准上评估SCDG:在PAN 2025衍生的成对基准上,SCDG达到0.92的精确率、0.97的召回率和0.94的F1值,优于所有基线;在PAN 2026的多源检索任务中,它达到0.83的nDCG@10和0.96的Recall@100,超过所有基线。在同主题、同事件的Multi-News测试中,经过校准的增益分布SCDG分类器仅对0.125%的对预测源复用,支持在该评估协议下对主题重叠的鲁棒性。这些结果确立了SCDG作为统一且可词分解的信号,适用于大量转换下的源特定内容复用。
英文摘要
Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexplored and largely unresolved challenge. Prior work on LLM-generated-text detection targets AI involvement, which may be permissible, rather than source reuse, while similarity-based methods struggle after extensive rewriting and multi-source synthesis. Motivated by the description-length view of probabilistic prediction, in which relevant side information can reduce a target sequence's code length, we introduce Source-Conditioned Description-Length Gain (SCDG), a directional, training-free framework that contrasts a frozen language model's description length of a suspicious document $P$ with and without a candidate source $S$. This contrast yields token-level log-likelihood gains that measure the incremental predictive evidence supplied by $S$. We evaluate SCDG on the PAN at CLEF benchmarks for generative plagiarism. On a PAN 2025-derived pairwise benchmark, SCDG achieves 0.92 Precision, 0.97 Recall, and 0.94 F1, outperforming all baselines; on PAN 2026's multi-source retrieval task, it reaches 0.83 nDCG@10 and 0.96 Recall@100, surpassing all baselines. On a same-topic, same-event Multi-News test, the calibrated gain-distribution SCDG classifier predicts source reuse for only $0.125\%$ of pairs, supporting robustness to topical overlap under this evaluation protocol. These results establish SCDG as a unified and token-decomposable signal for source-specific content reuse under extensive transformation.