arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ElecSafety:诊断大型语言模型对微控制器板的电气安全判断

ElecSafety: Diagnosing Large Language Model Safety Judgments for Microcontroller Boards

Linjian Yang, Xinyan Wang, Kunpeng Liu

arXiv 2610.06905首次发表:更新:

发表机构

Clemson University; Portland State University(克莱姆森大学; 波特兰州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM在微控制器板电气安全判断上的不足,提出ElecSafety基准测试,从厂商手册提取规则评估六种模型,发现规则可提升准确率但近边界场景仍不可靠。

AI 中文摘要

大型语言模型(LLMs)逐渐支持嵌入式硬件开发,但在作用于物理硬件时,一个看似合理的建议可能会带来安全风险。由于每块微控制器板的电气安全要求可能不同,因此产生了一个实际问题:LLMs能否具备对提议的用户操作是否对电路板安全做出正确决策的能力?尽管存在针对嵌入式开发的基准测试,但它们主要针对代码生成和硬件设计任务。这凸显了缺乏经过测试的、针对电路板特定电气安全判断的问题。为弥补这一空白,我们引入了ElecSafety基准测试,以评估LLM能否给出正确的标签,并从正确的制造商约束和规则中推导出该标签。我们从供应商数据手册和硬件手册中提取这些约束。每个场景由一个决定性条件和黄金标签(安全、危险或无法确定)定义,然后将场景转化为自然语言。我们将场景分为违规和合规案例、换板、近边界和缺失信息四类。随后,我们在六种开放权重LLM上,在规则访问和令牌限制的不同条件下评估了这些场景。研究发现,电路板规则可以提高准确率,而限制令牌的影响较小。即使是最佳配置,在近边界和缺失信息场景下仍然不可靠,且正确的标签往往依赖于不准确的约束。

英文摘要

Large language models (LLMs) progressively support embedded hardware development, but a plausible recommendation may pose safety risks when acting on the physical hardware. Because electrical safety can be different between every microcontroller board, a practical question is raised: can LLMs have the capability to make the right decision on whether a proposed user operation is safe for the board? Although benchmarks for embedded development exist, they primarily target code generation and hardware design tasks. This highlights the lack of tested board-specific judgment regarding electrical safety. To address this gap, we introduce the ElecSafety benchmark to evaluate whether an LLM can respond with the correct label and derive that label from the proper manufacturer constraints and rules. We extract these constraints from vendor datasheets and hardware manuals. Each scenario is defined by a decisive condition and a gold label (safe, hazardous, or cannot determine) before translating the scenario into natural language. We categorize the scenarios into violation and compliant cases, board-swap, near-boundary, and missing-information. We then evaluated the scenario on six open-weight LLMs under various conditions of rule access and token limitations. The finding shows board rules can improve accuracy, whereas limiting the token has a smaller impact. Even the best configuration remains unreliable on near-boundary and missing-information scenarios, and a correct label often rests on the inaccurate constraints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑