arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

签到式密钥检测:字符串即你所需

Checked-In Secret Detection: Strings Are All You Need

Zhengdong Huang, Kevin Li, Jinqiu Yang, Yepang Liu, Lili Wei

arXiv 2608.04523首次发表:更新:

AI 中文总结

该研究针对源代码硬编码密钥检测的现有方法局限,提出基于StringGroup的上下文提取算法及Secretron工具,在SecretBench数据集上F1值达98.74%,可高效检测未知密钥。

AI 中文摘要

源代码中硬编码的密钥会带来可被恶意攻击者轻易利用的严重安全漏洞。现有的基于正则表达式的检测方法存在根本性局限,因为密钥往往缺乏可识别的模式,导致精确率和召回率较低。近期研究探索了上下文感知的检测方法,因为周围代码可揭示候选字符串的用途,但这些方法面临三个关键挑战:(1)混淆鲁棒性问题,模型过度依赖易被混淆的标识符;(2)跨语言泛化困难,源于训练数据分布不均;(3)冗长且含噪声的上下文会引入过多无关标记并减慢推理速度。我们观察到字符串是代码语义的关键信息源,具有更高的上下文密度、混淆鲁棒性和语言独立性。基于此见解,我们提出StringGroup,一种挖掘潜在密钥周围字符串的新型上下文提取算法。通过对现有模式进行相对简单的修改,将分析范围专门缩小到字符串字面量,该方法取得了显著提升:仅使用原始上下文的33.2%,就能保留超过80%的语义信息,并大幅提高密钥检测的信噪比。我们进一步基于StringGroup方法和Transformer模型设计了上下文感知密钥检测工具Secretron。在SecretBench数据集上的评估显示,该工具准确率高,F1值达98.74%,在混淆和跨语言场景下具有强鲁棒性,优于最先进的基于大语言模型(LLM)的基线方法。我们将该工具部署到实际环境中,从26个应用中成功检测出48个此前未知的密钥,证明了我们方法的实际有效性。

英文摘要

Hardcoded secrets in source code pose critical security vulnerabilities which can be easily exploited by malicious adversaries. Existing regex-based detection approaches suffer from fundamental limitations, as secrets often lack identifiable patterns, resulting in poor precision and recall. Recent studies have explored context-aware detection methods, as surrounding code can reveal the purpose of candidate strings. However, these methods confront three key challenges: (1) obfuscation robustness where models over-rely on easily obfuscated identifiers, (2) cross-language generalization difficulties due to uneven training data distribution, and (3) lengthy and noisy context that introduces excessive irrelevant tokens and slows inference. We observe that strings serve as a critical information source for code semantics, offering superior contextual density, obfuscation robustness, and language independence. Based on this insight, we propose StringGroup, a novel context extraction algorithm that mines strings surrounding potential secrets. By introducing a relatively simple modification to existing patterns that narrows the analysis specifically to string literals, the method achieves significant gains. With only 33.2% of the original context, it preserves over 80% of semantic information and significantly improves the signal-to-noise ratio for secret detection. We further design a context-aware secret detection tool, Secretron, based on StringGroup methods and Transformer model. Evaluation on the SecretBench dataset demonstrates high accuracy with 98.74% F1-score and strong robustness under obfuscation and cross-language scenarios, outperforming state-of-the-art LLM-based baselines. We deploy our tool in real-world environments and successfully detect 48 previously unknown secret keys from 26 applications, demonstrating the practical effectiveness of our approach.

Journal refThe ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA) 2026

DOI:10.1145/3832207

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑