AI 中文总结
该研究提出结合多模态大语言模型智能体的主动学习框架,从ASAS-SN的37万余颗变星中高效发现异常天体,仅用约3小时、约40美元成本便得到含24个新异常的星表,验证了该方法的可行性。
AI 中文摘要
异常的光变曲线形态可能指向罕见的物理结构或新现象,但自动异常搜索常受伪影主导。将真实异常与误报分离传统上需要人工审查,这无法适配现代巡天的规模。我们提出一种用于检测周期性变星样本中异常的主动学习框架,将审查工作委托给多模态大语言模型智能体。初始排序基于在相位折叠光变曲线的DINOv2 ViT-g/14嵌入上训练的孤立森林算法得出。智能体会迭代审查排名最高候选者的光变曲线图像并分配相关性分数,该分数会通过嵌入空间传播以优先处理下一个目标。随后的逻辑回归步骤将搜索扩展至标签传播之外,多智能体共识审查则用于过滤误报。我们使用Gemini 3 Flash智能体,将该流程应用于ASAS-SN Sky Patrol V2.0中的373646颗周期性变星。经过51次迭代,智能体在约3小时内以约40美元的成本标记了约1%的样本,最终得到包含24个异常天体(其中18个为新报告)和153个潜在有趣天体的星表。这些异常包括具有极深主食的食双星、高振幅相接双星、脉动星以及爆发超过二十年的共生新星。在相同标记预算下,初始排序仅能恢复我们星表中30%的异常天体,而要找到排名最低的异常天体则需要约70倍的预算。这些结果证明了智能体主动学习在现有及未来测光巡天中用于异常检测的可行性。
英文摘要
Unusual light-curve morphologies can point to rare physical configurations or new phenomena, but automatic searches for anomalies are often dominated by artifacts. Separating genuine anomalies from false positives has traditionally required manual vetting, which does not scale to modern surveys. We present an active learning framework for detecting anomalies in samples of periodic variable stars, with the vetting delegated to multimodal large language model agents. The initial ranking comes from isolation forests trained on DINOv2 ViT-g/14 embeddings of phase-folded light curves. The agents iteratively review the light-curve images of the top-ranked candidates and assign relevance scores, which are propagated through the embedding space to prioritize the next targets. A logistic regression step then extends the search beyond label propagation, and a multi-agent consensus review filters false positives. Using Gemini~3 Flash agents, we applied the pipeline to $373\,646$ periodic variables from ASAS-SN Sky Patrol V2.0. Across $51$ iterations, the agents labeled ${\sim}1\%$ of the sample in roughly $3$ hours at a cost of ${\approx}$ $$40$, yielding our final catalog of $24$ anomalies ($18$ newly reported) and $153$ potentially interesting objects. The anomalies include eclipsing binaries with extremely deep primary eclipses, high-amplitude contact binaries and pulsators, and a symbiotic nova in outburst for over two decades. At the same labeling budget, the initial ranking would have recovered only $30\%$ of our catalog, while reaching its lowest-ranked anomaly would have required ${\sim}70$ times the budget. These results demonstrate the feasibility of agentic active learning for anomaly detection in existing and upcoming photometric surveys.
Comments24 pages, 12 figures, 6 tables. Submitted to ApJ