arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人类文本与AI生成文本的不可区分性研究

On the Indistinguishability of Human v/s AI Generated Text

Jaee Ponde, Aritra Das, Mihir More, Debayan Gupta

arXiv 2608.26797首次发表:更新:

发表机构

Truth Audit Labs(Truth Audit Labs)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLMs带来的AI与人类文本区分难题,提出利用人类写作样本改述机器生成文本的策略,推导收敛速率并分析有限样本下的相关缩放关系。

AI 中文摘要

大型语言模型(LLMs)的快速发展使得区分AI生成文本与人类写作成为一个紧迫问题,旨在让机器生成文本显得更“人性化”的改述工具进一步放大了这一挑战。我们研究如何利用人类写作样本,将机器生成的回应战略性地改述为贴近人类分布的文本。在针对相同提示词的人类与机器回应的多样本设置下,我们证明在简单混合与稳定性条件下,重复改述会使机器分布向经验人类分布移动。我们的结果推导出明确的收敛速率,将分析扩展到有限样本设置,并刻画所需人类样本数量与改述轮次如何随期望误差缩放。

英文摘要

The rapid improvement of LLMs has made distinguishing AI-generated text from human writing a pressing problem. This challenge is further amplified by paraphrasing tools designed to make machine-generated text appear more "human". We study how access to human writing samples can be used to strategically paraphrase machine-generated responses toward the human distribution. Under a multi-sample setting with human and machine responses to the same prompts, we show that repeated paraphrasing moves the machine distribution toward the empirical human distribution under simple mixing and stability conditions. Our results derive an explicit convergence rate, extend the analysis to a finite-sample setting, and characterize how the required number of human samples and paraphrasing rounds scale with the desired error.

Comments11 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑