arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向WASP-2025共享任务的基于SciBERT的高效上下文受限望远镜文献分类

Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT

Madhusudhana Naidu

arXiv 2609.01647首次发表:更新:

AI 中文总结

针对WASP-2025共享任务中望远镜文献分类的人工资源消耗问题,提出基于SciBERT的高效方法,在512token上下文限制下获0.89宏F1值,居任务排行榜榜首,为科学文本整理效率边界提供见解。

AI 中文摘要

望远镜文献的创建是评估天文台科学影响力、确保天文学研究可重复性的关键环节,该任务涉及识别、分类及关联引用或使用特定望远镜的科学出版物,但目前该过程仍以人工为主,资源消耗大。本研究提出一种基于SciBERT的高效方法,用于将科学论文自动分为四类:科学类、仪器类、提及类及非望远镜类。尽管存在严格的上下文长度限制(最大512个token)和有限的计算资源,该方法仍取得了0.89的宏F1值,在WASP-2025排行榜中位居榜首。研究分析了截断的影响,结果显示即便有一半样本超出token限制,SciBERT的领域适配性仍能实现稳健分类。此外,研究还探讨了截断、分块与长上下文模型之间的权衡,为科学文本整理的效率边界提供了见解。

英文摘要

The creation of telescope bibliographies is a crucial part of assessing the scientific impact of observatories and ensuring reproducibility in astronomy. This task involves identifying, categorizing, and linking scientific publications that reference or use specific telescopes. However, this process remains largely manual and resource intensive. In this work, we present an efficient SciBERT-based approach for automatic classification of scientific papers into four categories - science, instrumentation, mention, and not telescope. Despite strict context-length constraints (maximum 512 tokens) and limited compute resources, our approach achieved a macro F1 score of 0.89, ranking at the top of the WASP-2025 leaderboard. We analyze the effect of truncation and show that even with half the samples exceeding the token limit, SciBERT's domain alignment enables robust classification. We discuss trade-offs between truncation, chunking, and long-context models, providing insights into the efficiency frontier for scientific text curation.

Comments3 pages, 2 tables. 1st place system description for the TRACS shared task at WASP 2025 (Third Workshop for Artificial Intelligence for Scientific Publications), co-located with IJCNLP-AACL 2025. Published version: https://aclanthology.org/2025.wasp-main.21/ . Code: https://github.com/E0NIA/TRACS-WASP-2025-1st-Place

Journal refProceedings of the Third Workshop for Artificial Intelligence for Scientific Publications (WASP 2025), pages 192-194, Mumbai, India and virtual, December 2025. Association for Computational Linguistics

DOI:10.18653/v1/2025.wasp-main.21

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑