智能体时代的验证码:从每次遭遇中学习的求解器
CAPTCHAs in the Agentic Era: Solvers That Learn from Every Encounter
浏览论文内容
中文总结 AI 辅助
该研究提出的结合微调YOLOv8检测器与置信度路由开放权重VLM的自适应验证码求解器,可从遭遇中学习,在16类验证码上达85.4%整体准确率,能在对抗扰动中快速恢复,突破传统验证码防御的局限。
中文摘要 AI 辅助
视觉语言模型(VLMs)无需特定任务训练即可解决视觉验证码,但基于它们构建的智能体每次应对挑战都从零开始。对于此类智能体,一个熟悉谜题的第100个实例的耗时和计算量与第一个实例相同。专用检测器则权衡利弊,能在几毫秒内给出答案,但仅适用于其训练过的类别,且不会随着暴露于数据而改进。我们研究当求解器随着使用而改进时会发生什么。我们的系统将微调后的YOLOv8检测器与基于置信度路由的开放权重VLM配对,完全通过截图和操作系统输入事件运行,无需浏览器自动化或DOM访问。它在16个类别上达到整体准确率85.4%、宏观准确率84.2%,超过任一组件单独的性能。VLM生成的每个答案都可作为训练标签,因此检测器能吸收其从未训练过的类别,通常在一到两次遭遇后即可实现,且无需人工标注。相同的循环也能修复检测器:验证码操作员可对公开发布的检测器进行图像扰动,将其准确率降至0%,但这些扰动不会影响VLM,VLM的标签能让检测器恢复。在为期一年的模拟军备竞赛中,验证码操作员每月重新设计其扰动,求解器每一轮都能恢复,而一个准确率约70%的廉价开放权重教师能像完美的 oracle 一样有效地强化它。假设失败的机器人会一直失败的视觉验证码防御低估了自适应求解器恢复的速度。
英文摘要
Vision-language models (VLMs) can solve visual CAPTCHAs without task-specific training, but the agents built on them approach every challenge from scratch. For such an agent, the hundredth instance of a familiar puzzle costs as much time and compute as the first. Specialized detectors invert the trade-off, answering in milliseconds but only for categories they were trained on. Neither improves with exposure. We study what changes when a solver improves with use. Our system pairs a fine-tuned YOLOv8 detector with an open-weight VLM behind a confidence-based router, and runs entirely from screenshots and operating-system input events, with no browser automation or DOM access. It reaches 85.4% overall and 84.2% macro accuracy across 16 classes, exceeding either component alone. Every answer VLM produces also serves as a training label, so the detector absorbs categories it was never trained for, typically after one or two encounters and without human annotation. The same loop also repairs it. A CAPTCHA operator can perturb images against the publicly released detector and drive its accuracy to 0%, but the perturbations leave VLM untouched, and its labels let the detector recover. Under a year-long simulated arms race in which the CAPTCHA operator re-crafts its perturbations each month, the solver recovers every round, and a cheap ~70%-accurate open-weight teacher hardens it as effectively as a perfect oracle. Visual CAPTCHA defenses that assume a failing bot stays failing therefore understate how quickly an adaptive solver returns.
发表机构
- Istanbul Technical University(伊斯坦布尔理工大学)
机构由 AI 辅助整理,请以论文原文为准。