发表机构
University of Washington; Flatiron Institute(华盛顿大学; 平顿研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FOTO是面向天文学的图级语义搜索工具,利用开放嵌入模型和大型语言模型验证,在召回率上显著优于现有方法,并提升检索精度。
AI 中文摘要
天文学论文的许多内容以图表形式呈现,但文献检索索引的是文本,因此无法通过描述图表内容来查找图表。我们提出了FOTO,一个图级搜索工具,覆盖来自53,839篇同行评审的astro-ph论文的482,750条标题和说明文字。每个图表由其标题和说明文字索引一次,使用一个1.1亿参数的开放模型在本地运行进行嵌入,而大型语言模型仅应用于检索到的候选列表,用于验证真实图表,而非生成指向不存在图表的引用。在更严格的图级标准下,FOTO在召回率20时,根据语域不同,能恢复29%至79%的查询目标,而Pathfinder在较宽松的论文级标准下为6%至20%,Semantic Scholar为8%至16%,并且FOTO优于它所替代的付费嵌入API。对候选列表进行重排序后,在详细查询中,召回率@1从0.667提升至0.875;在模糊查询中,从0.071提升至0.571。开放的模型在验证任务上达到或超过付费的大型语言模型。
英文摘要
Astronomy papers carry much of their content in figures, but literature search indexes text, so there is no way to find a plot by describing what it shows. We present FOTO, a figure-level search tool over 482,750 captions from 53,839 peer-reviewed astro-ph papers. Each figure is indexed once by its title and caption, embedded with a 110M-parameter open model that runs locally, and a large language model is applied only to the retrieved shortlist, where it verifies real figures rather than generating references to ones that do not exist. Held to the harder figure-level criterion, FOTO recovers the target for 29 to 79% of queries at recall 20 depending on register, against 6 to 20% for Pathfinder and 8 to 16% for Semantic Scholar on the easier paper-level criterion, and it beats the paid embedding API it replaced. Reranking the shortlist then raises recall at 1 from 0.667 to 0.875 on detailed queries and from 0.071 to 0.571 on vague ones. Openly available models match or beat paid LLM models for verification