arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02983cs.CRcs.SE

基于模式的密钥检测的边界变异测试:一种规则级方法及跨扫描器评估

Boundary-Mutation Testing for Pattern-Based Secret Detection: A Rule-Level Method and Cross-Scanner Evaluation

Shweta Mishra

AI总结:

本研究提出用于密钥检测的边界变异测试方法,经跨扫描器评估发现多款密钥扫描器存在边界脆弱性等缺陷,验证了修复方案并揭示了真实代码与合成语料库的误报差异等结果。

AI中文摘要:

基于模式的密钥扫描器通常通过基于示例的测试用例进行验证,这类测试用例仅固定了一个变量:凭证周围的文本。我们引入边界变异测试来改变该上下文,从每个规则自身的正则表达式生成凭证,将其嵌入真实的源代码上下文,并在规则级别而非工具级别对结果进行分类,从而得到三个检测指标。将该方法应用于三个扫描器——包含43条规则的开源扫描器Gitleaks 8.21.2和TruffleHog 3.82.13,在10种上下文下,主要主体的检测率保持在≥0.9976,但当凭证以连字符结尾时,检测率骤降至0.5233。有5条规则受到影响:2条带有固定数量量词的规则完全且确定性失效;3条带有可变数量量词的规则发生回溯并匹配截断的凭证;熵 fallback(回退)机制挽救了部分失效,但降低了其严重程度。我们验证了一种修复方案,该方案可恢复全部鲁棒性且无新的误报,并指出一个需注意的问题:一次看似合理的首次尝试悄悄导致两条规则退化,仅通过重新运行同一测试集才被发现。Gitleaks存在一个与源相关的缺陷——硬编码的终止符允许列表导致大多数凭证类型完全漏检,而TruffleHog无边界脆弱性但覆盖范围最窄。我们报告的是边际而非条件失效概率:一种常见令牌格式在结构上免疫,另一种在64次中失效1次;在292527行真实代码上,误报排序与合成语料库相反。由于每种类型的检测是确定性的,比较单位是10种凭证类型而非数百个样本;无召回率差异达到显著性,因此我们报告零结果而非排名。

英文摘要:

Pattern-based secret scanners are commonly validated with example-based fixtures that fix one variable: the text surrounding a credential. We introduce boundary-mutation testing to vary that context, generating credentials from each rule's own regular expression, embedding them in realistic source contexts, and classifying outcomes at the rule level rather than the tool level, yielding three detection metrics. Applied to three scanners - a 43-rule open-source scanner, Gitleaks 8.21.2, and TruffleHog 3.82.13 - detection in the primary subject holds at >=0.9976 across ten contexts but collapses to 0.5233 when a credential ends in a hyphen. Five rules are affected: two, with fixed-count quantifiers, fail totally and deterministically; three, with variable-count quantifiers, backtrack and match a truncated credential; an entropy fallback rescues some failures but downgrades their severity. We validate a repair restoring full robustness with no new false positives, and report a caution: a plausible first attempt silently regressed two rules, caught only by re-running the same battery. Gitleaks has an unrelated, source-confirmed defect - a hard-coded terminator allowlist causing total misses for most credential types - while TruffleHog shows no boundary fragility but narrowest coverage. We report marginal, not conditional, failure probabilities: one common token format is structurally immune, another fails once in 64; on 292,527 lines of real code, the false-positive ordering inverts relative to the synthetic corpus. Because per-type detection is deterministic, the comparison unit is ten credential types, not hundreds of samples; no recall difference reaches significance, so we report the null result, not a ranking.

补充信息

↑