arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06037cs.IR

基于LLM的科学论文实体解析的视觉分析

Visual Analysis of LLM-based Entity Resolution from Scientific Papers

Siyu Wu, Yi Yang, Weize Wu, Ruiming Li, Yuyang Zhang, Ge Wang, Huobin Tan, Zipeng Liu, Lei Shi

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种结合大型语言模型与视觉分析的人机协作流程,用于从科学文献中批量提取特定领域实体,通过人在回路优化,将单文档实体解析准确率提升约30%。

中文摘要 AI 辅助

本文聚焦于从大量科学文献中提取特定领域实体的视觉分析支持,这一任务在使用传统命名实体解析方法时存在固有局限性。随着如GPT-4等大型语言模型(LLMs)的出现,由于LLM在实体解析中具备整合多种文本理解等能力,相较于传统机器学习方法取得了显著改进。本研究提出了一种新的视觉分析流程,将先进的LLMs与多功能的可视化和交互设计相结合,以支持批量实体解析。具体而言,我们关注金属有机框架(MOFs)这一特定材料科学领域,以及名为CSD-MOFs的大型数据集。通过与材料科学领域专家的合作,我们获得了标注良好的合成段落。我们提出了在实体解析过程中采用视觉分析技术进行人在回路(human-in-the-loop)的优化,使领域专家能够交互式地将见解融入LLM智能中,包括错误分析和检索增强生成(RAG)算法的解释。我们通过RAG示例选择案例研究的评估表明,这种人机协作方法将单文档实体解析准确率提高了约30%。

英文摘要

This paper focuses on the visual analytics support for extracting domain-specific entity from extensive scientific literature, a task with inherent limitations using traditional named entity resolution methods. With the advent of large language models (LLMs) such as GPT-4, significant improvements over conventional machine learning approaches have been achieved due to LLM's capability on entity resolution integrate abilities such as understanding multiple types of text. This research introduces a new visual analysis pipeline that integrates these advanced LLMs with versatile visualization and interaction designs to support batch entity resolution. Specifically, we focus on a specific material science field of Metal-Organic Frameworks (MOFs) and a large data collection namely CSD-MOFs. Through collaboration with domain experts in material science, we obtain well-labeled synthesis paragraphs. We propose human-in-the-loop refinement over the entity resolution process using visual analytics techniques, which allows domain experts to interactively integrate insights into LLM intelligence, including error analysis and interpretation of the retrieval-augmented generation (RAG) algorithm. Our evaluation through the case study of example selection for RAG demonstrates that this human-machine collaborative approach improved single-document entity resolution accuracy by approximately 30%.

发表机构

  • School of Software, Beihang University(北京航空航天大学软件学院)
  • School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
  • School of Materials Science and Engineering, University of Science and Technology Beijing(北京科技大学材料科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑