arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CESBench:面向物联网设备密码工程安全的大语言模型基准测试

CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices

Wenquan Zhou, An Wang, Jing Liang, Peien Feng, Jingqi Zhang, Yaoling Ding, Liehuang Zhu

arXiv 2609.21344首次发表:更新:

发表机构

Beijing Institute of Technology; Shandong University(北京理工大学; 山东大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CESBench基准,含380个专家条目覆盖物联网密码工程安全六子领域,评估11个LLM,发现安全判定理由生成是最弱能力。

AI 中文摘要

对于物联网(IoT)设备而言,仅靠安全的算法是不够的:拥有物理访问权限的攻击者可以直接攻击实现,而且一旦部署,其缺陷很难修复。大语言模型(LLMs)现在被用于构建和分析此类实现。现有的LLM基准测试涵盖密码学和通用网络安全,但均未覆盖密码工程。在本文中,我们提出CESBench,包含380个由专家编写的条目,覆盖物联网设备密码工程安全的六个子领域:侧信道、故障注入、实现、对策、评估和集成。四种任务类型针对不同能力:209个多项选择题测试记忆,67个判断题要求给出安全判定及其理由,63个场景题要求进行工程诊断,41个代码任务由572个测试用例评分。为验证基准测试,11个开放权重和专有LLM回答了每个条目。多项选择和代码响应自动评分,判断和场景响应由LLM评判员评分,其分数与来自另一模型族的第二个评判员及人工重新评分进行核对。综合得分范围为54.4%至83.6%。每种任务类型的最高分为:多项选择98.6%,代码95.1%,场景诊断88.4%,但判断仅为58.8%。在所有模型中,88.5%的判定正确,但其理由仅获得评分标准中53.4%的分数。对于最强模型,多项选择接近其上限,且大多数代码任务已解决,而证明安全判定的合理性仍是最弱的能力。基准测试、提示词和逐条目结果均已公开。

英文摘要

For Internet of Things (IoT) devices, a secure algorithm alone is not enough: an attacker with physical access can attack the implementation directly, and its flaws are hard to fix once deployed. Large language models (LLMs) are now used to build and analyze such implementations. LLM benchmarks exist for cryptography and general cybersecurity, but none covers cryptographic engineering. In this paper, we present CESBench, 380 expert-written items across six sub-domains of cryptographic engineering security for IoT devices: side-channel, fault injection, implementation, countermeasures, evaluation, and integration. Four task types target different competences: 209 multiple-choice items test recall, 67 judgment items require a security verdict and its justification, 63 scenario items require an engineering diagnosis, and 41 code tasks are graded by 572 test cases. To validate the benchmark, 11 open-weight and proprietary LLMs answer every item. Multiple-choice and code responses are scored automatically, and judgment and scenario responses by an LLM judge, whose scores are checked against a second judge from another model family and human re-scoring. Composite scores range from 54.4% to 83.6%. The top score on each task type is 98.6% for multiple choice, 95.1% for code, and 88.4% for scenario diagnosis, but only 58.8% for judgment. Across models, 88.5% of verdicts are correct, yet their justifications earn only 53.4% of the rubric marks. Multiple choice is near its ceiling for the strongest models and most code tasks are solved, whereas justifying a security verdict remains the weakest competence. The benchmark, prompts, and per-item results are public.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑