arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06960cs.SEcs.CR

面向物联网漏洞检测的反编译可执行性与程序等价性统计分析

Statistical Analysis of Executability and Program Equivalence in Decompilation for IoT Vulnerability Detection

Minami Yoda, Jialong Li, Yasuyuki Tahara, Yuichi Sei, Yutaka Matsuno

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对物联网固件漏洞检测的反编译问题,提出九维质量评估指标,通过统计分析证明其可有效评估反编译质量,为物联网设备缺陷及黑盒生成模型质量评估提供基础。

中文摘要 AI 辅助

物联网(IoT)设备处理用户音频、视频、认证数据等敏感隐私信息,因此检测其固件漏洞至关重要。反编译是一项关键检测技术,近年来因大语言模型(LLM)具备高可读性和高重编译成功率而受到关注。但LLM输出依赖概率性token预测,倾向于优先保证语法正确性,可能生成看似合理却与原始二进制语义不同的代码,漏洞常出现在错误处理流程、边界检查等易在此过程中丢失的细节中。现有评估指标主要关注测试用例通过情况,无法充分识别出内部结构被改变却行为看似有效的代码,因此需要从多个角度量化反编译代码内部结构的指标。本文提出一个包含结构、行为、语义相似度三类的九维质量评估指标。针对开源路由器平台OpenWrt上的318个程序,使用五种方法(一种基于规则、四种基于LLM)生成19625个反编译结果并进行统计分析。重编译成功组的总体得分显著高于失败组(Cohen's d=0.92);行为相似度的d值为0.96,结构相似度的d值为0.69,表明这些指标是反编译质量的重要预测因子。本研究为量化物联网设备的实现缺陷提供了统计评估基础,也为黑盒生成模型的质量评估提供了可推广的框架。

英文摘要

Internet of Things (IoT) devices handle sensitive privacy-related information such as user audio, video, and authentication data, making it essential to detect vulnerabilities in their firmware. Decompilation, a key detection technique, has recently attracted attention because Large Language Models (LLMs) enable high readability and high recompilation success rates. However, because LLM outputs depend on probabilistic token prediction, they tend to prioritize syntactic correctness and may generate plausible-looking code that is semantically different from the original binary. Vulnerabilities often arise in details that are easily lost in this process, such as error-handling flows and boundary checks. Existing evaluation metrics focus mainly on passing test cases and cannot sufficiently identify code whose internal structure has been altered despite appearing behaviorally valid, so a metric that quantifies the internal structure of decompiled code from multiple perspectives is needed. We propose a nine-dimensional quality evaluation metric consisting of three categories: structural, behavioral, and semantic similarity. Targeting 318 programs from OpenWrt, an open-source router platform underlying many commercial routers, we generated 19,625 decompilation results using five methods (one rule-based and four LLM-based) and analyzed them statistically. The recompilation-success group achieved significantly higher overall scores than the failure group (Cohen's d=0.92); behavioral similarity showed d=0.96 and structural similarity d=0.69, demonstrating that these metrics are important predictors of decompilation quality. This study provides a statistical evaluation foundation for quantifying implementation defects in IoT devices and a framework that generalizes to quality evaluation of black-box generative models.

补充信息

↑