D2C-Routing:用于混合来源AI生成文本检测的维度到组成证据路由
D2C-Routing: Dimension-to-Composition Evidence Routing for Mixed-Origin AI-Generated Text Detection
浏览论文内容
中文总结 AI 辅助
本研究针对混合来源AI生成文本检测问题,提出D2C-Routing方法,在MixD2C数据集上取得0.8603的四向平均TPR@1%FPR,性能优于RACE-local重运行结果。
中文摘要 AI 辅助
AI生成文本检测通常被视为判断文本是人类撰写还是机器生成的二元文档级任务,但这种框架在混合来源写作中失效,这类写作的内容来源与表达来源可能不同。我们将混合来源检测视为维度到组成的来源归因,先推断内容来源与表达来源,再将其组合为四种协作类型。我们提出维度到组成路由(D2C-Routing),它将内容侧与表达侧证据路由至有监督维度头,再通过学习的门控组合层预测最终标签。在从HART混合来源基准重构的划分MixD2C上,我们公开的基于D2C-Routing的检测器系统达到0.8603的四向平均真阳性率@1%假阳性率,比同一划分的RACE-local重运行结果高6.5个点。核心 ablation 实验支持该路由设计,误差分析显示,区分AI内容/人类表达与完全AI生成文本仍是最困难的边界。代码可在该https URL获取。
英文摘要
AI-generated text detection is commonly framed as a binary document-level judgment about whether a text is human-written or machine-generated. This framing breaks down for mixed-origin writing, where content origin and expression origin may differ. We cast mixed-origin detection as dimension-to-composition source attribution, inferring content origin and expression origin before composing them into four collaboration types. We propose Dimension-to-Composition Routing (D2C-Routing), which routes content-side and expression-side evidence to supervised dimension heads before a learned gated composition layer predicts the final label. On MixD2C, a reconstructed split derived from the HART mixed-origin benchmark, our disclosed D2C-Routing-based detector system reaches 0.8603 four-way Avg TPR@1%FPR, 6.5 points above the same-split RACE-local rerun. Core ablations support the routing design, while error analysis shows that distinguishing AI-content/human-expression from fully AI-generated text remains the hardest boundary. Code is available at https://github.com/bystander563/d2c-routing-artifact.
发表机构
- Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。