arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26893cs.CR

ACTS:在受控盲测条件下评估LLM密码识别能力的多层级基准

ACTS: A multi-tier benchmark evaluating LLM cipher identification under controlled blind conditions

Youssef Hamdi Zafaan Ibrahim, Mohammed Khalaf Salama

首次发表
浏览论文内容

中文总结 AI 辅助

ACTS基准通过分层元数据剥夺和强制推理测试,评估LLM在盲测条件下的密码识别能力,发现其准确率仅略高于随机基线,且存在显著的元数据依赖和扩展失败问题。

中文摘要 AI 辅助

我们提出了ACTS(密码测试套件中的伪影),一个可复现的基准测试,通过分层元数据剥夺(第1层:完整元数据;第2层:仅文件名;第3层:完全盲测)来隔离密码分析能力,并仅基于密文测试强制推理(第5层:思维链、代码推理、自我修正)。一项在7,000个文件上进行的10配置消融研究,使用单一70/30训练-测试划分进行特征移除分析,提供了大规模额外证据。在v2b语料库上进行实时API推理,使用完全随机的填充(每个文件选择PKCS7、ISO 10126和ANSI X9.23)、唯一的CSPRNG密钥和唯一明文(每个模型127个文件,跨三个云系统的381条第1层记录和380条第3层记录,并辅以第4层中254次自主智能体评估),得出合并的第3层准确率为30.8%,仅略高于七类分类的14.3%随机基线。相应的合并元数据依赖差距为40.9个百分点(第1层:71.7%对比第3层:30.8%)。与经典随机森林(在7,000个文件上,基于工程化字节级特征训练,准确率为69.2%)相比,观察到的实时差距为38.4个百分点。由于此比较跨越不同的输入表示和训练范式,该差距应被解释为整体能力差异,而非纯粹的因子分解。报告了六项发现:(1)元数据依赖性仍然很大;(2)盲测条件下的扩展失败;(3)强制推理是附带现象;(4)观察到的实时能力差距超过早期启发式估计;(5)信号主要是结构性的而非统计性的;(6)启发式不变性与机器学习脆弱性揭示了不同的失败模式。

英文摘要

We introduce ACTS (Artifacts in Cipher Testing Suite), a reproducible benchmark that isolates cryptanalytic ability through tiered metadata deprivation (Tier-1: full metadata; Tier-2: filename only; Tier-3: completely blind) and tests forced reasoning (Tier-5: chain-of-thought, code-as-reasoning, self-correction) on ciphertext alone. A 10-configuration ablation study on 7,000 files, using a single 70/30 train-test split for feature-removal analysis, provides additional evidence at scale. Live API inference on a v2b corpus with fully randomised padding (PKCS7, ISO 10126, and ANSI X9.23 selected per file), unique CSPRNG keys, and unique plaintexts (127 files per model, 381 Tier-1 records, 380 Tier-3 records across three cloud systems, supplemented by 254 autonomous agentic evaluations in Tier-4) yields a combined Tier-3 accuracy of 30.8%, only modestly above the 14.3% random baseline for seven-way classification. The corresponding combined metadata-dependency gap is 40.9 percentage points (Tier-1: 71.7% vs. Tier-3: 30.8%). Against a classical Random Forest (69.2% on 7,000 files, trained on engineered byte-level features), the observed live gap is 38.4 percentage points. Because this comparison spans different input representations and training paradigms, the gap should be interpreted as an overall capability difference rather than a clean factorial decomposition. Six findings are reported: (1) Metadata dependency remains large; (2) Scaling failure under blind conditions; (3) Forced reasoning is epiphenomenal; (4) The observed live capability gap exceeds earlier heuristic estimates; (5) The signal is primarily structural rather than statistical; (6) Heuristic invariance versus ML fragility reveals different failure modes.

补充信息

↑