arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于预算自适应信号分词的指令条件电磁频谱理解

Instruction-Conditioned Electromagnetic Spectrum Understanding via Budget-Adaptive Signal Tokenization

Lei Zhai, Zhihao Chang, Shuyuan Yang, Zhixi Feng

arXiv 2610.12142首次发表:更新:

发表机构

School of Artificial Intelligence, Xidian University(西安电子科技大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对电磁频谱监测需求,提出预算自适应信号分词器BATok与多模态指令数据集EMSpec-Instruct,实现将VLMs扩展至I/Q原始信号,在相关任务中取得了有竞争力的性能。

AI 中文摘要

电磁频谱监测日益需要超越特定任务识别与检测的灵活分析。多模态大语言模型提供了统一接口,但将视觉语言模型(VLMs)扩展至同相/正交(I/Q)原始信号需在保真度与严格预算间权衡的分词方案。对于信号,密集编码会使分词成本随观测长度增长,而固定分辨率压缩可能丢失短时长或局部信号证据。因此,我们提出BATok,一种预算自适应信号分词器,可根据输入长度调整分词容量,并依据信号内容分配该容量。BATok通过轻量多分辨率分支从信号衍生特征构建候选表示,再结合局部能量先验与可学习查询,将这些表示重采样为紧凑信号分词。分词数量适配输入长度,同时保持严格边界。生成的分词被投影至VLMs的语言嵌入空间。我们进一步引入EMSpec-Instruct,一种多模态指令数据集,对齐I/Q原始信号、瀑布图图像与语言监督,用于调制识别、结构化检测及语言条件信号 grounding。实验表明,BATok学习到有效的信号表示,并在所有任务中取得了有竞争力的性能。

英文摘要

Electromagnetic spectrum monitoring increasingly requires flexible analysis beyond task-specific recognition and detection. Multimodal large language models offer a unified interface, but extending vision-language models (VLMs) to raw I/Q signals requires tokenization that balances fidelity against a strict budget. For signals, dense encoding causes token costs to grow with observation length, whereas fixed-resolution compression may discard short-duration or localized signal evidence. Thus, we propose \textbf{BATok}, a budget-adaptive signal tokenizer that adjusts token capacity to the input length while allocating that capacity according to the signal content. BATok constructs candidate representations from signal-derived features using lightweight multi-resolution branches, then combines a local energy prior with learnable queries to resample these representations into compact signal tokens. The number of tokens adapts to the input length while remaining strictly bounded. The resulting tokens are projected into the language embedding space of VLMs. We further introduce \textbf{EMSpec-Instruct}, a multimodal instruction dataset aligning raw I/Q signals, waterfall images, and language supervision for modulation recognition, structured detection, and language-conditioned signal grounding. Experiments show that BATok learns effective signal representations and achieves competitive performance across all tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑