arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以毒攻毒:利用对抗机器学习保护练习题抵御AI作弊的可行性研究

Fighting Fire with Fire: On the Feasibility of Protecting Exercises Against AI Cheating

Tobias Braun, Jonas Grebe, Louis Rethfeld, Marcus Rohrbach

arXiv 2608.01112首次发表:更新:

AI 中文总结

本研究探究利用对抗机器学习,通过在多模态选择题视觉组件添加扰动引导AI作弊者给出固定错误答案,再以统计检验检测作弊,以保护教育练习题抵御AI作弊的可行性。

AI 中文摘要

生成式AI的广泛应用使学生能够将认知任务外包给日益强大的助手,造成能力假象的同时,破坏了教育旨在培养的独立推理能力。本研究探究是否可将对抗机器学习重新用于保护教育练习题,抵御此类有害依赖。所提方法采用多模态选择题,其视觉组件可通过细微视觉扰动进行保护,这些扰动会引导AI求解器偏向指定的错误答案。这些响应形成统计指纹:盲目复制求解器答案的学生,会比真正自主作答的学生更频繁地重现诱导出的答案模式。在现实黑盒助手假设下,研究人员使用三种最常见的前沿多模态语言模型——Anthropic的Claude、Google的Gemini和OpenAI的ChatGPT,验证该范式的可行性。通过使用可访问的替代模型,研究人员优化了能诱导一致响应模式的对抗扰动,这些模式可通过统计假设检验实现有依据的检测。这些发现确立了用机器自身漏洞对抗机器辅助推理的潜力与局限性。

英文摘要

The widespread adoption of generative AI enables students to outsource cognitive effort to increasingly capable assistants, creating an illusion of competence while undermining the independent reasoning that education aims to cultivate. We investigate whether adversarial machine learning can be repurposed to protect educational exercises against such corrosive reliance. Our approach uses multimodal multiple-choice questions whose visual components can be protected with subtle visual perturbations that steer AI solvers toward designated incorrect answers. These responses form a statistical fingerprint: students who blindly copy a solver reproduce the induced answer pattern more frequently than genuine students. We study the feasibility of this paradigm under realistic black-box assistant assumptions using three of the most common state-of-the-art multimodal language models: Anthropic's Claude, Google's Gemini, and OpenAI's ChatGPT. By using accessible surrogate models, we optimize adversarial perturbations that induce consistent response patterns. Those patterns enable principled detection through statistical hypothesis testing. These findings establish both the promise and the limitations of fighting machine-assisted reasoning with the vulnerabilities of the machines themselves.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑