发表机构
University of Macau; SRH University; Populus Group(澳门大学; SRH大学; 波普勒斯集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出针对印地语、泰卢固语和泰米尔语的基准IndicDetect,评估AI生成文本检测器在分布偏移等现实场景下的鲁棒性,发现现有检测器鲁棒性不足,为相关检测任务奠定基础。
AI 中文摘要
大语言模型(LLM)的快速普及进一步提升了开发可靠的AI生成文本检测技术的需求,尤其是针对英语之外的语言。然而,当前的基准测试对印度诸语言关注甚少,且在无法代表真实世界的理想化环境中测试检测器。我们提出了一个针对印地语、泰卢固语和泰米尔语的AI生成文本检测通用基准,命名为IndicDetect,旨在评估检测器在现实分布偏移下的鲁棒性。IndicDetect包含精心筛选的人类撰写文本,以及来自不同领域和生成器的LLM生成对应文本,并在存在领域偏移、生成器偏移和对抗扰动的情况下系统地评估检测器。使用单一且可重复的评估方案,我们评估了多种统计型和神经型检测器。我们发现存在显著的鲁棒性失效:监督式神经检测器在分布内表现良好,而无训练方法在未见生成器和对抗攻击下性能大幅下降。这些失效的严重程度因语言而异,其中印地语在对抗扰动下的整体性能下降最为显著。这些结果表明,现有检测器在印度诸语言场景中的主要弱点在于鲁棒性,而非峰值准确率。IndicDetect提供了标准数据划分、评估协议和基线,为印度文字的AI生成文本检测建立了鲁棒且感知语言的基础。
英文摘要
The rapid proliferation of LLMs has further heightened the need to develop dependable AI-generated text detection, especially beyond English. Nevertheless, current benchmarks pay little attention to Indic languages and test detectors in idealized settings that do not represent the real world. We present a generalized benchmark for AI-generated text detection in Hindi, Telugu, and Tamil, which we call IndicDetect, designed to assess the robustness of detectors under realistic distribution shifts. IndicDetect comprises highly curated human-written texts matched with LLM-generated counterparts across various domains and generators, and systematically evaluates detectors in the presence of domain shift, generator shift, and adversarial perturbation. Using a single and repeatable evaluation scheme, we evaluate a wide range of statistical and neural detectors. We find substantial robustness failures: supervised neural detectors perform well in-distribution, while training-free methods degrade considerably under unseen generators and adversarial attacks. The severity of these failures varies across languages, with Hindi exhibiting the largest overall degradation under adversarial perturbations. These results highlight that the primary weakness of existing detectors in Indic settings lies in their robustness, not in their peak accuracy. IndicDetect provides standard data splits, an evaluation protocol, and baselines to establish a robust, language-aware foundation for AI-generated text detection in Indic scripts.