AI 中文总结
针对现有MGT检测仅评估全生成文本、忽视人机协作的局限,提出中文基准C-HAT-Bench,涵盖5,000源文本与240,000变体,评测21个检测器,发现协作模式下AUROC平均下降12.0%,跨模式迁移不对称。
AI 中文摘要
大型语言模型(LLMs)越来越多地通过修改或扩展人类草稿参与写作,导致机器参与在形式和程度上均有所变化。然而,大多数机器生成文本(MGT)检测器仅在完全人类撰写与完全AI生成文本的二元设置下进行评估。由于人机协作可能削弱或重新分配与机器生成相关的线索,在这种二元设置下的强性能可能高估检测器的可靠性。这一不匹配在中文语境中仍未得到充分探索:检测线索受分词和语言特定文本分布的影响,但涵盖生产设置、领域和生成器的受控资源仍然有限。为填补这一空白,我们提出了中文人机协作文本检测基准(C-HAT-Bench),这是一个统一基准,将来自五个领域的5,000篇人类撰写源文本与使用六种生成模型在前缀条件续写作为参考设置及三种协作生产模式下产生的超过240,000个变体相连接。我们通过四种协议评估了21个检测器,涵盖零样本和预训练监督的文档级检测、边界定位以及跨条件泛化。相对于前缀条件续写,文档级检测器在协作生产模式下的平均AUROC降低了12.0%,其中最大的检测器特定相对降幅达到44.4%。跨协作生产模式的迁移也是不对称的,表明在给定生产设置中的性能并不能可靠预测其他生产设置中的性能。
英文摘要
Large Language Models (LLMs) increasingly participate in writing by modifying or extending human drafts, causing machine involvement to vary in both form and extent. Yet most Machine-Generated Text (MGT) detectors are evaluated only on fully human-written versus fully AI-generated text. Because human--AI collaboration can weaken or redistribute cues associated with machine generation, strong performance under this binary setting may overstate detector reliability. This mismatch remains underexplored in Chinese: detection cues are shaped by tokenization and language-specific text distributions, yet controlled resources spanning production settings, domains, and generators remain limited. To fill this gap, we present a Chinese Human-AI Collaborative Text Detection Benchmark (C-HAT-Bench), a unified benchmark that links $5,000$ human-written source texts from five domains to more than $240,000$ variants produced using six generative models under Prefix-Conditioned Continuation as a reference setting and three collaborative production modes. We evaluate $21$ detectors through four protocols spanning zero-shot and pretrained supervised document-level detection, boundary localization, and cross-condition generalization. Relative to Prefix-Conditioned Continuation, mean AUROC across document-level detectors is $12.0\%$ lower on the collaborative production modes, with the largest detector-specific relative decrease reaching $44.4\%$. Transfer across collaborative production modes is also asymmetric, indicating that performance in a given production setting is not a reliable predictor of performance in other production settings.