发表机构
University of Colorado Denver(科罗拉多大学丹佛分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TCA-SIR将科学灵感检索重新定义为面向目标的抽象学习,在ResearchBench上较MOOSE-Chem提升HitRate@top4%超10个百分点,实现更优检索与可解释性。
AI 中文摘要
用于科学的人工智能(AI for Science)的科学假设生成通常包括科学灵感检索(SIR),随后是假设构建。现有的SIR方法按主题相似性对论文进行排名,未明确表示候选灵感如何迁移到目标问题,这对于远程灵感尤其受限,其价值往往在于可复用的解决问题原则而非主题重叠。受人类抽象源的可迁移方面并将其重新映射到新目标的方式启发,我们将SIR重新表述为面向目标的抽象(TCA):检索对象是从候选中专门为目标提取的可迁移抽象原则。我们提出TCA-SIR,该模型学习生成面向目标的抽象表示,并使用其表示来预测可迁移性。在ResearchBench上,TCA-SIR的性能优于现有SIR方法和直接大语言模型(LLM)检索,其HitRate@top4%较MOOSE-Chem提升超过10个百分点;学习到的抽象表示也比未训练的TCA提示更清晰地恢复与目标相关的机制,既实现了更强的检索效果,又为科学灵感提供了可解释的依据。
英文摘要
Scientific hypothesis generation for AI for Science typically involves Scientific Inspiration Retrieval (SIR) followed by hypothesis composition. Existing SIR methods rank papers by topical similarity and do not explicitly represent how a candidate inspiration transfers to a target problem. This is especially limiting for remote inspirations, whose value often lies in reusable problem-solving principles rather than topical overlap. Motivated by how humans abstract transferable aspects of a source and remap them to a new target, we reformulate SIR as target-conditioned abstraction (TCA). The retrieval object is a transferable abstract principle extracted from a candidate specifically for the target. We present TCA-SIR, which learns to generate target-conditioned abstractions and uses their representations to predict transferability. On ResearchBench, TCA-SIR outperforms prior SIR methods and direct LLM retrieval, improving HitRate@top4% over MOOSE-Chem by more than 10 percentage points. Learned abstractions also recover target-relevant mechanisms more clearly than an untrained TCA prompt, yielding both stronger retrieval and an interpretable rationale for scientific inspiration.