arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Writerslogic 在 PAN 2026:域偏移下基于过程而非内容的稳健检测

Writerslogic at PAN 2026: Process over Content for Robust Detection under Domain Shift

David L. Condrey

arXiv 2610.03565首次发表:更新:

发表机构

WritersLogic Inc(WritersLogic 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一个基于过程而非内容的特征稳健性框架,用于域偏移下的检测任务,并在 PAN 2026 三个任务中取得领先成绩,验证了域不变特征的有效性。

AI 中文摘要

我们描述了 Writerslogic 系统参与 PAN at CLEF 2026 的三个共享任务(推理轨迹检测、Voight-Kampff 生成式 AI 检测、多作者写作风格分析),这些任务由一个共享的分析框架统一:在分布偏移下,特征的稳健性由训练分布与测试分布之间的支持重叠决定,而非由训练集效应量决定。这产生了一个分类法(域锚定、域可移植、域不变),解释了为什么生成器特定特征在域偏移下失效,而词汇指纹(hapax 比率、Yule's K、Heaps 指数)、压缩度量以及字符 n-gram 能够存活。在推理轨迹检测任务中,训练完全基于数学,而 84% 的测试数据来自未见过的域,该框架指导系统设计在源检测中获得第一名(通过 Opus-Sonnet 一致性达到 0.85 宏 F1),在安全分类中获得第三名(通过查询-拒绝分解达到 0.66 宏 F1)。对于 Voight-Kampff 任务,我们构建了一个校准集成,包括 DeBERTa-v2(ONNX)、具有 44 个域可移植文体特征的多种子 LightGBM,以及基于 n-gram TF-IDF 的 SVM,通过学习的堆叠和等渗校准进行组合;最佳配置在 PAN 2026 测试集上达到 0.891,各评估维度的平衡子指标在 0.853 到 0.902 之间。对于多作者写作风格分析,我们描述了一个系统,融合了基于字符 n-gram 相似性图的谱聚类、用于局部边界检测的归一化压缩距离,以及用于神经变点检测的 SmolLM-135M 困惑度;由于平台混淆,我们的运行从未到达官方评估,因此我们报告了设计及其先验预测。在所有三个任务中,测量生成过程属性的特征被设计为在域偏移下优于测量生成内容属性的特征。

英文摘要

We describe the Writerslogic systems for three PAN at CLEF 2026 shared tasks (Reasoning Trajectory Detection, Voight-Kampff Generative AI Detection, and Multi-Author Writing Style Analysis), unified by a shared analytical framework: feature robustness under distribution shift is governed by support overlap between training and test distributions, not by training-set effect size. This yields a taxonomy (domain-anchored, domain-portable, domain-invariant) that explains why generator-specific features die under domain shift while vocabulary fingerprints (hapax ratio, Yule's K, Heaps' exponent), compression measures, and character n-grams survive. On Reasoning Trajectory Detection, where training was entirely mathematics and 84 percent of test was unseen domains, the framework guided system design to 1st place in source detection (0.85 macro F1 via Opus-Sonnet agreement) and 3rd place in safety classification (0.66 macro F1 via query-refusal decomposition). For Voight-Kampff, we built a calibrated ensemble of DeBERTa-v2 (ONNX), multi-seed LightGBM with 44 domain-portable stylometric features, and SVM on n-gram TF-IDF, combined via learned stacking with isotonic calibration; the best configuration achieved 0.891 on the PAN 2026 test set with balanced sub-metrics (0.853 to 0.902 across all evaluation dimensions). For Multi-Author Writing Style Analysis, we describe a system fusing spectral clustering over character n-gram similarity graphs, normalized compression distance for local boundary detection, and SmolLM-135M perplexity for neural change-point detection; a platform mix-up meant our run never reached the official evaluation, so we report the design and its a priori predictions. Across all three tasks, features measuring generation process properties are designed to outperform features measuring generated content properties under domain shift.

Comments13 pages, 1 figure, 6 tables. Notebook for the PAN Lab at CLEF 2026. Code https://github.com/dcondrey/voight-kampff-clef2026 and https://github.com/dcondrey/trajectory-detection-clef2026

Journal refCLEF 2026 Working Notes, CEUR Workshop Proceedings, pp. 5474-5486

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑