arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17644astro-ph.IMcs.AIcs.CYcs.LG

重新思考天文学语言模型在开放式科学推理中的领域专业化

Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models

  • Oak Ridge National Laboratory(橡树岭国家实验室)
  • The Ohio State University(俄亥俄州立大学)
  • QUP/IPNS, High Energy Accelerator Research Organization (KEK)(高能加速器研究机构(KEK)QUP/IPNS)
  • Max-Planck-Institut für Astronomie(马克斯·普朗克天文研究所)

机构由 AI 辅助整理,请以论文原文为准。

Vanessa Lama, Sanjay Das, Emily Herron, Yuan-Sen Ting, Tijmen de Haan, Junqi Yin, Tirthankar Ghosal, Feiyi Wang

AI总结:

本研究通过天文学问答基准比较通用与领域专门化模型,发现通用模型正确性更高,主张领域专业化应视任务和部署而定,并强调领域特定评估的重要性。

AI中文摘要:

领域专门化的语言模型被广泛用于科学问答,但更强大的通用系统提出了一个更尖锐的问题:在开放式科学推理中,领域特定的微调何时仍然有价值?我们通过一个基于公开可得的2017--2026年奥林匹克风格材料构建的精选问答基准,在天文学领域对此进行了研究。自由回答子集包含300个问题,其中204个为纯文本问题,96个为图像关联问题。我们使用基于评判者的正确性和互补的参考指标,比较了开放权重和API服务的通用、多模态以及天文学专门化模型。强大的通用模型在该测试平台上建立了最高的正确性基线,而指标一致性、评判者敏感性、基准构成和模态的分析揭示了单一排行榜无法捕捉的差异。这些结果促使我们将领域专业化视为一个任务和部署相关的属性,并强调了领域特定评估在确定哪些模型、能力和评估标准适用于科学工作流程中的作用。

英文摘要:

Domain-specialized language models are widely used for scientific question answering, but stronger general-purpose systems raise a sharper question: when does domain-specific fine-tuning remain valuable for open-ended scientific reasoning? We study this in astronomy with a curated QA benchmark from publicly available 2017--2026 Olympiad-style materials. The free-response subset contains 300 questions, including 204 text-only and 96 image-linked examples. We compare open-weight and API-served general-purpose, multimodal, and astronomy-specialized models using judge-based correctness and complementary reference metrics. Strong general-purpose models establish the highest correctness baseline in this testbed, while analyses of metric agreement, judge sensitivity, benchmark composition, and modality reveal variation not captured by a single leaderboard. These results motivate treating domain specialization as a task- and deployment-dependent property and highlight the role of domain-specific evaluation in determining which models, capabilities, and evaluation criteria are appropriate for scientific workflows.

↑