发表机构
Universidade Federal de Campina Grande; Instituto de Formação de Educadores, Universidade Federal do Cariri(坎皮纳格兰德联邦大学; 卡里里联邦大学教育工作者培训学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究通用视觉语言模型在零样本下检测快速射电暴,从模拟动态频谱中抽取样本作基准,比较其与专用探测器,发现Gemma 4 2B准确率与SwinYNet相近,在结构化RFI上误报率低,重写提示可用于三类分类,准确率较高。
AI 中文摘要
快速射电暴(FRBs)是持续时间为毫秒级的射电瞬变现象,其自动检测越来越依赖高度专业化的深度学习模型。这些探测器性能卓越,但需要大量特定任务训练数据集且重新定义需重新训练。本文评估小型、开放权重、可本地运行的通用视觉语言模型(VLMs)能否在零样本、仅提示的情况下检测动态频谱中的FRBs,无需微调与标记示例,并给出结构化决策及自然语言理由。从3000个包含FRBs、结构化射频干扰(RFI)和噪声的模拟L波段动态频谱中抽取2000个样本的平衡二元基准,逐样本比较两个此类VLMs(Gemma 4 2B和4B)与最先进的专用探测器SwinYNet。默认阈值下,Gemma 4 2B准确率达93.65%,与SwinYNet(92.90%)无显著差异,在结构化RFI上误报率更低(6.4%对25.0%)且纯噪声无误报。SwinYNet在此基准上保持完美概率排名(ROC-AUC为1.0000对0.9482),零样本VLM仅通过通用预训练接近此上限。仅重写提示就能在3000个频谱全集上为三类FRB/RFI/噪声分类重新配置相同模型,准确率高达86%且无FRB误报。
英文摘要
Fast Radio Burst (FRB) detection increasingly relies on specialized deep learning models that require large task-specific training sets and cannot be redefined without retraining. We evaluate whether small, open-weight, locally run generalist Vision-Language Models (VLMs) can detect FRBs in dynamic spectra under a zero-shot, prompt-only regime. On a balanced binary benchmark of 2000 simulated L-band spectra, Gemma 4 E2B reaches an accuracy of 94.05\%, statistically indistinguishable from the specialized detector SwinYNet (92.85\%), with a far lower false-positive rate on structured RFI (4.8\% vs. 24.6\%) and none on pure noise, though SwinYNet ranks perfectly (ROC-AUC 1.0000 vs. 0.9520). Rewriting the prompt alone reconfigures the same models for three-class FRB/RFI/noise classification, reaching up to 86.0\% accuracy without a single false FRB while classifying each 2 s spectrum in 1.0--1.5 s, faster than the observation itself. Applied unchanged to the 1600 real FAST observations of FAST-FREX, they reject real interference almost perfectly (2 and 5 false positives in 1000 negatives) but recover only 28.5\% and 27.0\% of the 600 catalogued bursts, against 95.7\% reported for SwinYNet on the same files. Stratifying those bursts by the dispersed signal in the image shows the limit to be the input representation rather than the classifier, recall rising to 84--85\% where the sweep is unambiguous and collapsing to 1\% on the 13\% of positives carrying no detectable signal in a 2 s undedispersed full-band view. The simulated bursts are nearly 30 times brighter in median, and at matched brightness the recalls agree to within a few points.
Comments31 pages, 7 figures. Section added with analysis for real data. New figures and tables added. Other minor changes