Doc2LoRA 提供科学思想的可解码表示
Doc2LoRA Provides Decodable Representations of Scientific Ideas
- Binghamton University(宾汉姆顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出Doc2LoRA方法,用LoRA适配器表示论文,使向量空间中的点可解码为语言模型,同时支持搜索与生成,并在APS数据集上验证了其有效性。
AI中文摘要:
将科学论文表示为空间中的点,使我们能够搜索相似论文,并探究各领域之间的相互关系及创新驱动因素。除搜索之外,论文向量空间还支持生成:通过简单的向量运算混合论文可创建新的点,这反映了组合式新颖性,即将现有思想重新组合为新思想。然而,混合点往往代表尚未有任何论文实现的思想,附近也没有论文可识别该思想。我们提出通过 Doc-to-LoRA 超网络生成的 LoRA 适配器来表示每篇论文。因此,空间中的每个点(包括混合点)都代表一个对自然语言中的问题和指令开放的大型语言模型(LLM)。在美国物理学会(APS)的论文上,我们指示每个子领域平均点处的 LLM 用几个词命名该领域,得到的标签比五个基线的标签更接近官方名称,这通过词重叠和由五个 LLM 评审员组成的小组进行评判。我们还要求位于两篇 APS 论文之间的点处的 LLM 撰写摘要,得到的描述随混合权重的变化从一篇论文转向另一篇论文。虽然 Doc-to-LoRA 是为生成而训练的,但一个小的可逆变换使嵌入在搜索方面具有竞争力,与 SPECTER2 和 EmbeddingGemma 相当,并接近 SBERT。由于该变换是可逆的,变换后空间中的每个点仍可映射回一个 LLM。因此,嵌入同时服务于搜索和生成,使研究人员能够质疑空间中任何点的思想,并将其作为生成新思想的起点。
英文摘要:
Representing scientific papers as points in a space lets us search for similar papers and inquire about how fields relate to one another and drive innovation. Beyond search, the vector space of papers invites generation: mixing papers through simple vector operations creates new points, mirroring combinatorial novelty, the recombination of existing ideas into new ones. However, a mixed point often represents an idea no paper has yet realized, with no papers nearby to identify the idea. We propose representing each paper by a LoRA adapter generated by the Doc-to-LoRA hypernetwork. Every point in the space, including mixtures, thus represents a large language model (LLM) open to questions and instructions in natural language. On papers from the American Physical Society (APS), we instruct the LLM at the average of each subfield to name the field in a few words and obtain labels closer to the official names than the labels of five baselines, as judged by word overlap and a panel of five LLM judges. We also ask the LLMs at points between two APS papers to write an abstract and obtain descriptions shifting from one paper to the other in step with the mixing weight. While Doc-to-LoRA is trained for generation, a small invertible transform makes the embeddings competitive for search, on par with SPECTER2 and EmbeddingGemma and close to SBERT. Because the transform is invertible, every point in the transformed space still maps back to an LLM. The embeddings thus serve both search and generation, enabling researchers to question the idea at any point in the space as a starting point for generating new ideas.