AI 中文总结
本文提出领域感知RAG流水线,结合8种模型与4种硬件兼容性验证器,检测嵌入式C软件函数复用,解决SonarQube高假阳性问题,验证器准确率达97.5%。
AI 中文摘要
在不同产品间复用嵌入式软件函数具有经济价值,但技术难度大:即便两个函数的余弦相似度得分高于0.90且均通过SonarQube质量检查,针对不同微控制器平台实现的相同功能在硬件层面仍可能完全不兼容。静态分析工具旨在测量代码质量而非硬件域兼容性,不具备外设接口、硬件抽象层(HAL)依赖或寄存器映射约束的模型。本文提出一种面向嵌入式C软件复用检测的领域感知检索增强生成(RAG)流水线,直接解决硬件兼容性缺口。该流水线通过提取每个函数的现有内联注释、调用图上下文和项目README对其进行丰富,再以8种骨干模型(MiniLM、MPNet、BGE、E5、GraphCodeBERT、OpenAI text-embedding-3-small、LLaMA 3 8B、StarCoder2 3B)作为特征提取器进行嵌入。4种硬件兼容性验证器(覆盖外设令牌重叠、参数数量奇偶性、调用图依赖重叠及结构分支模式)直接在检索栈中过滤候选对象。在6个公开嵌入式C软件项目(184个函数、4815个平台间配对)上评估,该流水线显示SonarQube作为复用过滤器的假阳性率达93.6%,其中83.5%的失败由静态分析无法检测的硬件环境不匹配导致。对40个被拒绝配对的手动验证确认验证器准确率为97.5%,诊断规则注入变体识别出主要失败类别(McNemar卡方≈294.0,p<0.001)。
英文摘要
Reusing embedded software functions across products is economically valuable but technically difficult: the same functionality implemented for two different microcontroller platforms can be entirely incompatible at the hardware level, even when the functions score above 0.90 cosine similarity and both pass SonarQube quality checks. Static analysis tools were designed to measure code quality, not hardware-domain compatibility, and have no model of peripheral interfaces, hardware abstraction layer (HAL) dependencies, or register-map constraints. This paper presents a domain-aware retrieval-augmented generation (RAG) pipeline for embedded C software reuse detection that addresses the hardware-compatibility gap directly. The pipeline enriches each function by extracting its existing inline comments, call-graph context, and a project README before embedding it with eight backbone models (MiniLM, MPNet, BGE, E5, GraphCodeBERT, OpenAI text-embedding-3-small, LLaMA 3 8B, StarCoder2 3B) acting as feature extractors. Four hardware-compatibility validators---covering peripheral token overlap, parameter count parity, call-graph dependency overlap, and structural branching pattern---filter candidates directly in the retrieval stack. Evaluated on six public embedded C software projects (184 functions, 4,815 above-plateau pairs), the pipeline reveals that SonarQube produces a 93.6% false-positive rate as a reuse filter, with 83.5% of failures caused by hardware-environment mismatches that static analysis cannot detect. Manual verification of 40 rejected pairs confirms 97.5% validator accuracy, and a diagnostic rule-injection variant identifies the dominant failure categories (McNemar chi-squared~=~294.0, p~$<$~0.001).