AI 中文总结
研究从非结构化文档提取秘密及访问上下文的问题,提出多智能体大语言模型系统SSA,它能通过检测与审查智能体配对提取相关信息,经合成基准评估,相比传统方法提升了提取精度、召回率,速度更快,让凭证检测更具可操作性。
AI 中文摘要
暴露的文档(如电子邮件、聊天记录、工单和事件记录)经常泄露凭证,但在事件响应中,泄露的秘密只是问题的一半。响应者还需要识别秘密打开的“门”:凭证可能允许攻击者访问的账户、租户、端点、数据库、云资源或其他系统。传统的秘密扫描器依赖正则表达式或训练好的分类器,在格式良好的代码上效果良好,但在凭证碎片化、重新格式化或远离其解锁的资源时会遇到困难,且只报告秘密字符串而不说明其打开的对象。我们提出了秘密扫描代理(SSA),这是一个多智能体大语言模型系统,能从非结构化暴露文档中提取秘密及其相关的“门”以及支持证据。SSA将一个注重召回率的检测智能体与一个过滤误报并恢复缺失上下文的审查智能体配对。由于真实凭证数据敏感,我们在自己生成的涵盖23种秘密类型和多种文档格式的合成基准上评估SSA,通过编程匹配、大语言模型判断和人工审查的三步流程进行评分。在六个模型中,多智能体SSA比单智能体变体提高了提取精度,在“门”提取方面提升最大,高达16个百分点。SSA在匹配正则表达式扫描器精度的同时,召回率提高了两倍多,与13名安全分析师相比,它更精确,能恢复近两倍的秘密-“门”对,且运行速度快5到17倍。通过在一个结果中返回秘密、其“门”和支持证据,SSA将凭证检测转化为可用于分类和补救的可操作发现。
英文摘要
Exposed documents such as emails, chat threads, tickets, and incident notes routinely leak credentials, but during incident response a leaked secret is only half the story. Responders also need to identify the ``door'' the secret opens: the account, tenant, endpoint, database, cloud resource, or other system that the credential could allow an attacker to access. Traditional secret scanners rely on regular expressions or trained classifiers which work well on well-formatted code, yet they struggle when a credential is fragmented, reformatted, or far from the resource it unlocks, and they report the secret string without naming what it opens. We present Secret Scanner Agent (SSA), a multi-agent large-language-model system that extracts both the secret and its associated door, together with supporting evidence, from unstructured exposed documents. SSA pairs a detection agent that favors recall with a review agent that filters false positives and recovers missing context. Because real credential data is sensitive, we evaluate SSA on synthetic benchmarks we generated that span 23 secret types and multiple document formats, scored with a three-step pipeline of programmatic matching, an LLM judge, and human review. Across six models, multi-agent SSA improves extraction precision over a single-agent variant, with the largest gains on door extraction, by up to 16 percentage points. SSA matches a regular-expression scanner's precision while more than tripling its recall, and against thirteen security analysts it is more precise, recovers nearly twice as many secret--door pairs, and runs five to seventeen times faster. By returning the secret, its door, and supporting evidence in one result, SSA turns credential detection into an actionable finding for triage and remediation.
CommentsSubmitted to the Conference on Applied Machine Learning for Information Security (CAMLIS) 2026