发表机构
School of Cyber Science and Engineering, Huazhong University of Science and Technology; Hubei Key Laboratory of Distributed System Security; University of Georgia(华中科技大学网络空间安全学院; 湖北省分布式系统安全重点实验室; 佐治亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究基于大规模MCU固件数据集,实证评估24种仿真固件分析工具,发现仅34.5%样本可模糊测试且覆盖率低,揭示了现有方法的局限并指明未来研究方向。
AI 中文摘要
基于微控制器(MCU)的设备日益普及,使得高效、可扩展的固件安全分析变得至关重要。最近的固件重托管工作实现了自动化漏洞评估,但仍存在两个空白。首先,现有工具在有限且高度重叠的数据集上进行评估,削弱了所报告结果的普适性。其次,先前研究沿孤立维度推进了技术水平,例如仿真、模糊测试或漏洞诊断,而没有整体理解这些组件在动态分析工作流中如何相互作用和互补。基于最近从OTACAP和FirmLine发布的大规模MCU固件数据集,我们对发表在顶级会议和期刊上的24种基于仿真的固件分析工具进行了实证研究。我们将这些工具定位在一个统一的自动化分析流程中:仿真配置侦察、仿真、漏洞发现和诊断。在每个阶段,我们使用去重且验证过的4,571个基于ARM的固件样本子集,评估工具的输出是否提供足够信息以支持后续阶段。即使将一个样本视为成功(如果至少有一个现有工具可以对其进行模糊测试),也只有1,580个样本(34.5%)可以成功进行模糊测试。其中,模糊测试结果普遍较差,平均代码覆盖率仅为10%,且存在许多误报的崩溃/挂起。通过对失败的模糊测试尝试和误报案例的系统分析,我们揭示了当前方法中的基本挑战和方法论局限性。这些发现突出了现有技术中的关键空白,并为指导未来基于仿真的固件分析研究提供了可行的见解。
英文摘要
Microcontroller (MCU)-based devices are increasingly pervasive, making efficient, scalable firmware security analysis critical. Recent firmware re-hosting work enables automated vulnerability assessment, yet two gaps remain. First, existing tools are evaluated on limited, heavily overlapping datasets, undermining the generalizability of reported results. Second, prior research advances the state of the art along isolated dimensions, such as emulation, fuzzing, or bug diagnosis, without a holistic understanding of how these components interact and complement one another in dynamic analysis workflows. Building on recently released large-scale MCU firmware datasets from OTACAP and FirmLine, we present an empirical study of 24 emulation-based firmware analysis tools published in top conferences and journals. We position these tools within a unified automated analysis pipeline: emulation configuration reconnaissance, emulation, bug finding, and diagnosis. At each stage, we evaluate whether a tool's output provides sufficient information to enable the subsequent stage, using a deduplicated and validated subset of 4,571 ARM-based firmware samples. Only 1,580 samples (34.5%) can be successfully fuzzed, even counting a sample successful if at least one existing tool can fuzz it. Among these, fuzzing results are generally poor, with average code coverage of only 10% and many false crashes/hangs. Through a systematic analysis of failed fuzzing attempts and false-positive cases, we expose fundamental challenges and methodological limitations in current approaches. These findings highlight critical gaps in the state of the art and provide actionable insights to guide future research in emulation-based firmware analysis.
CommentsExtended Version of NDSS 2027 Accept Paper