智能体可利用基础模型规避AI检测
Agents Can Use Base Models to Evade AI Detection
AI总结:
本研究展示编码智能体通过拼接基础模型样本来规避AI文本检测,降低检测率并保持任务性能,但成本显著增加。
AI中文摘要:
我们证明,配备基础语言模型的编码智能体能够通过拼接其样本中的内容来成功组装响应,从而规避检测。基础模型已被证明能够规避商业检测器,然而,先前的“人性化”技术依赖于使用这些模型对AI输出进行多轮迭代改写,这不可避免地导致语义漂移。相比之下,为编码智能体配备直接编排写作过程的能力,通过拼接来自基础模型的文本样本,使其能够生成连贯、任务特定且总体高质量的输出。我们发现,在Claude Code框架中运行的Claude Opus 5能够有效编排一个本地的32B参数OLMo-2基础语言模型,并在涵盖创意写作、事实依据、健康问答和指令遵循的基准测试中几乎不损失任务准确性,同时使用高达90%的基础模型令牌。以这种方式构建的响应降低了事后检测器(Pangram v4检测率从77%降至24%)和预先应用于智能体生成的软水印(在低误报率下模拟检测率降至10%)的有效性。尽管有效,但这种规避需要智能体显著更多的输入和输出令牌,在API定价下,每次查询的美元成本增加高达30倍。总体而言,这项工作展示了一类针对AI文本检测的新型对抗攻击的有效性,并敦促事后检测提供商在其训练中纳入基础模型的输出。
英文摘要:
We show that coding agents equipped with a base language model can successfully assemble responses from its samples to evade detection. Base models have been shown to evade commercial detectors, however, prior "humanization" techniques rely on using these models to paraphrase AI outputs over several iterations, which invariably results in semantic drift. In contrast, equipping coding agents to directly orchestrate the writing process by stitching text samples from a base model allows it to produce outputs that are coherent, task-specific and generally high quality. We find that Claude Opus 5 operating in a Claude Code harness effectively orchestrates a local 32B parameter OLMo-2 base LM and sacrifices little task accuracy across benchmarks spanning creative writing, factual grounding, health QA and instruction following, while using up to 90% base LM tokens. Responses constructed in this manner reduce the effectiveness of both post-hoc detectors (Pangram v4 detection rate drops from 77% to 24%) and soft watermarking applied a priori to the agent's generations (down to a simulated 10% detection at low FPR). While effective, this evasion requires a significantly larger number of input and output tokens from the agent, increasing the dollar cost per query up to 30x at API-pricing. Overall, this work demonstrates the effectiveness of a new class of adversarial attacks against AI text detection, and urges post-hoc detection providers to include outputs of base models in their training.