arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向嵌入式软件开发的大语言模型智能体的闭环评估

Closed-loop evaluation of LLM agents for embedded software development

Jorge García-Carrasco, Sergio García-Carrasco, Alejandro Maté, Juan Trujillo

arXiv 2610.11447首次发表:更新:

发表机构

Lucentia Research, Department of Software and Computing Systems, University of Alicante; Universitary Institute of Materials Technology (IUTM), Universitat Politècnica de València(阿利坎特大学软件与计算系统系卢森西亚研究; 巴伦西亚理工大学材料技术大学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出嵌入式编码智能体闭环评估基准,评估7种GPT与Qwen系列模型,发现gpt-5.4通过率最高,qwen3.5-27B为最强本地模型,表明高性能本地嵌入式编码智能体正在兴起。

AI 中文摘要

大语言模型(LLM)正越来越多地被用作编码智能体,用于编辑文件、运行构建和测试、检查执行结果并迭代修复软件。嵌入式固件是一个要求严苛的目标,因为其正确性取决于感知、时序和安全约束下的闭环行为,而非仅静态源码质量。然而,嵌入式智能体的评估仍有限,且常强调单次合成或离线正确性。我们提出了一个用于嵌入式编码智能体闭环评估的基准。每个任务提供纯文本工程描述、受限工作空间以及可见的构建与运行时界面。智能体必须将需求转化为实现和自验证步骤,然后迭代直至达到所需设备行为。该套件包含5个嵌入式控制任务和4种反馈场景:单次生成、真实自验证、CI风格的红/绿反馈以及神谕风格的详细反馈。实现针对模拟ESP32固件以确保可复现性。我们在5个任务和4种场景下评估了7种GPT系列和Qwen系列配置,每个条件重复3次,共420次运行。gpt-5.4在评估配置中通过率最高,但未达到基准饱和;qwen3.5-27B是观测到的最强本地模型;较小的本地模型在通过率和搜索效率上大幅下降。这些结果表明,具备能力的本地嵌入式编码智能体正在出现。

英文摘要

Large language models (LLMs) are increasingly deployed as coding agents that edit files, run builds and tests, inspect execution results, and repair software iteratively. Embedded firmware is a demanding target because correctness depends on closed-loop behavior under sensing, timing, and safety constraints, not only on static source quality. Yet embedded-agent evaluation remains limited and often emphasizes one-shot synthesis or offline correctness. We present a benchmark for closed-loop evaluation of embedded coding agents. Each task provides a plain-text engineering description, constrained workspace, and visible build-and-runtime surface. The agent must translate requirements into implementation and self-verification steps, then iterate until the required device behavior is achieved. The suite contains five embedded-control tasks and four feedback scenarios: one-shot generation, realistic self-verification, CI-style red/green feedback, and oracle-style detailed feedback. The implementation targets simulated ESP32 firmware for reproducibility. We evaluate seven GPT-family and Qwen-family configurations across five tasks and four scenarios, with three repetitions per condition for 420 runs. gpt-5.4 has the highest pass rate among evaluated configurations but does not saturate the benchmark; qwen3.5-27B is the strongest observed local model; and smaller local models degrade sharply in pass rate and search efficiency. These results suggest that capable local embedded coding agents are emerging.

CommentsPublished in Journal of Systems Architecture 179 (2026) 103937. 23 pages, 4 figures, 5 tables. Code and artifacts: https://github.com/jgcarrasco/closed_loop_evaluation_agents_embedded

Journal refJ. Syst. Archit. 179 (2026) 103937

DOI:10.1016/j.sysarc.2026.103937

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑