arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24086cs.AIcs.CEcs.SE

EMRB:用于评估大型语言模型对原始电磁信号推理能力的多级基准

EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • The 63rd Research Institute, National University of Defense Technology(国防科技大学第六十三研究所)
  • Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

Mingxu Zhang, Ying Sun, Yuhan Li, Yang Ji, Dazhong Shen, Ke Zhang, Shan Huang

AI总结:

本研究提出EMRB基准评估LLMs对原始电磁信号的推理能力,针对14个LLMs的实验显示其表现存在差距,提出的ReconPilot方法可显著提升LLMs的相关任务得分。

AI中文摘要:

大型语言模型(LLMs)越来越多地被用作科学与工程分析的代码智能体,但其分析原始物理层测量数据的能力尚未得到验证。我们推出EMRB(Electromagnetic Reasoning Benchmark,电磁推理基准),该基准通过编写和运行代码来评估LLMs分析原始I/Q数据的能力。EMRB包含200个问题,分为5个难度级别和27种问题类型,涵盖从信号检测到OFDM设计的任务,由11种带有验证真值的信号类型生成。与基于预处理特征或结构化表格构建的基准不同,EMRB仅提供原始采集数据,每个问题所涉及的量必须先通过代码发现。我们评估了14个LLMs,涵盖专有、开放权重和推理导向的系列,其得分范围为24.1%至78.9%,其中基本测量任务的平均得分从84.9%降至系统设计任务的21.2%。我们还提出了ReconPilot,一种将信号侦察、针对性分析和自我验证分离的结构化方法。在三种主干模型上,ReconPilot使总体得分提高了3.8至17.6个百分点,并改善了所测试的15种主干模型组合中的13种。所有数据和代码均在我们的GitHub代码库中公开提供。

英文摘要:

Large language models (LLMs) are increasingly used as code agents for scientific and engineering analysis, but their ability to analyze raw physical-layer measurements remains untested. We introduce \textbf{EMRB} (\textbf{E}lectro\textbf{m}agnetic \textbf{R}easoning \textbf{B}enchmark), which evaluates whether LLMs can analyze raw I/Q data by writing and running code. EMRB contains 200 problems across five difficulty levels and 27 question types, from signal detection to OFDM design, generated from 11 signal types with verified ground truth. Unlike benchmarks built on preprocessed features or structured tables, EMRB provides only the raw capture; the quantities each question refers to must first be discovered through code. We evaluate 14 LLMs spanning proprietary, open-weight, and reasoning-oriented families. Scores range from 24.1\% to 78.9\%, with the mean dropping from 84.9\% on basic measurement to 21.2\% on system design. We also propose \textbf{ReconPilot}, a structured method that separates signal reconnaissance, targeted analysis, and self-verification. Across three backbones, ReconPilot raises the overall score by 3.8 to 17.6 points and improves 13 of 15 backbone-level combinations tested. All data and code are publicly released in \href{https://github.com/mingxuZhang2/EMRB}{\textcolor{blue}{our GitHub repository}}.

↑