用通用大语言模型破解暗网验证码
Breaking Darknet CAPTCHAs with general purpose LLM
浏览论文内容
中文总结 AI 辅助
本研究针对暗网验证码,提出结合MLLM与经典计算机视觉的混合框架,通过任务重构或专用工具缓解MLLM的几何处理局限,实现90%以上的破解成功率。
中文摘要 AI 辅助
本研究评估了破解暗网环境中常见验证码挑战的自动化方法的有效性。这类验证码通常设计为不依赖JavaScript运行,因此与主流验证码系统相比具有独特特征。研究考虑了三种代表性挑战类型:开环定位、基于旋转的对齐以及对象选择验证码。实验揭示了当代多模态大语言模型(MLLM)的系统性局限:尽管它们通常能够识别相关视觉结构,但在精确空间定位和几何变换方面经常存在困难。这些缺陷可通过任务重构或为模型配备专用图像处理工具来缓解。因此,我们提出一种混合框架,其中MLLM作为高层推理与协调层,通过模型上下文协议(MCP)将几何计算委托给确定性算法。该系统在所有评估的验证码类型中均达到90%以上的成功率,表明结合MLLM与经典计算机视觉的互补优势,可得到比单独使用任一方法更准确高效的求解器。
英文摘要
Our work evaluates the effectiveness of automated methods for solving CAPTCHA challenges commonly encountered in darknet environments. These CAPTCHAs are typically designed to operate without JavaScript, resulting in distinct characteristics compared to mainstream CAPTCHA systems. Our study considers three representative challenge types: open-circle localization, rotation-based alignment, and object-selection CAPTCHAs. The experiments reveal a systematic limitation of contemporary MLLMs: while they are generally capable of identifying relevant visual structures, they frequently struggle with precise spatial localization and geometric transformations. These deficiencies can be mitigated either through task reformulation or by augmenting the models with specialized image processing tools. These deficiencies can be mitigated by task reformulation or by equipping the model with specialized image-processing tools. We therefore propose a hybrid framework in which an MLLM serves as a high-level reasoning and orchestration layer while delegating geometric computations to deterministic algorithms via the Model Context Protocol (MCP). The resulting system achieves success rates above 90% across all evaluated CAPTCHA types and demonstrates that combining the complementary strengths of MLLMs and classical computer vision yields a more accurate and efficient solver than either approach alone.
发表机构
- Bern University of Applied Science (BFH)(伯尔尼应用科学大学(BFH))
机构由 AI 辅助整理,请以论文原文为准。