语言模型能否理解毫米波数据?基于毫米波雷达的人类理解任务对大语言模型进行基准测试
Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding
- DGIST(大邱庆北科学技术院)
- KAIST(韩国科学技术院)
- InnoCORE LLM
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究针对毫米波数据与大语言模型融合的研究空白,构建首个毫米波人类感知基准mmWave-QA,评估LLMs在该任务上的零样本推理潜力与视觉退化下的鲁棒性,为相关研究奠定基础。
中文摘要 AI 辅助
大语言模型(LLMs)展现出卓越的推理与生成能力,推动其被用作感知领域的通用推理引擎。尽管视觉-语言模型(VLMs)等现代方法已尝试将推理能力融入视觉感知,但LLMs与毫米波(mmWave)模态的融合——尽管毫米波在低光照和遮挡场景下具有独特优势——仍未得到充分探索。主要瓶颈在于雷达-语言配对数据稀缺、跨数据集异质性严重,以及缺乏基础的毫米波编码器。我们通过一个极简的文本化接口解决这一问题,该接口将每个毫米波点云序列化为简洁的自然语言,使现成的LLMs可在问答(QA)场景中运行。在此基础上,我们推出mmWave-QA,这是首个针对语言条件下毫米波人类感知的基准测试。mmWave-QA整合了异质性的公开毫米波数据集,通过校准感知预处理和全局分类法对齐对其进行统一,同时提供自然语言问答任务。该基准测试涵盖6种场景和5项问答任务,支持在不同毫米波硬件和实验条件下进行标准化评估,为毫米波-LLM融合的可扩展研究奠定了基础。我们还在mmWave-QA上对LLMs进行评估与分析,重点关注它们在雷达感知方面的零样本推理潜力,以及在视觉退化场景下的鲁棒性。
英文摘要
Large language models (LLMs) have shown remarkable reasoning and generative capabilities, motivating their use as universal reasoning engines for perception. While modern approaches such as vision-language models (VLMs) have attempted to incorporate reasoning capabilities into visual sensing, the integration of LLMs with the millimeter-wave (mmWave) modality-despite its unique advantages under low light and occlusion-remains largely unexplored. The principal bottlenecks stem from the scarcity of radar language pairs, severe cross-dataset heterogeneity, and the absence of a foundational mmWave encoder. We address this gap through a minimal textualization interface that serializes each mmWave point cloud into concise natural language, allowing off-the-shelf LLMs to operate in a question answering (QA) setting. Building on this, we present mmWave-QA, the first benchmark for language-conditioned mmWave human perception. mmWave-QA aggregates heterogeneous public mmWave datasets and harmonizes them via calibration-aware preprocessing and global taxonomy alignment, while providing natural language QA. Spanning six scenarios and five QA tasks, the benchmark enables standardized evaluation across diverse mmWave hardware and experimental conditions, establishing a foundation for scalable research on mmWave-LLM integration. We further evaluate and analyze LLMs on our mmWave-QA, highlighting their zero-shot reasoning potential for radar perception, as well as their robustness under visual degradation.