arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SALSA:半自主文献摘要助手

SALSA: Semi-Autonomous Literature Summarization Assistant

William Schertzer, Sonakshi Gupta, Rampi Ramprasad

arXiv 2609.22210首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SALSA是一个开源的人机协同平台,结合文档解析、大语言模型、OCR、计算机视觉等技术,从多模态文献中自动提取结构化科学数据,支持可定制工作流,旨在实现可扩展、可靠的数据整理,适用于材料研究等领域。

AI 中文摘要

SALSA(半自主文献摘要助手)是一个开源的、人在回路中的平台,用于从多模态文献来源中提取结构化科学数据集。该软件结合了文档解析、大型语言模型、光学字符识别、计算机视觉、图形数字化和用户引导的校正工具,以从文本、表格、图形和标题中恢复结构化信息。用户可以配置提取阶段、定义数据集模式、对数字化图形进行干预,并导出经过验证的数据以供下游分析和机器学习使用。SALSA旨在自动化重复性的文献整理任务,同时在需要专家判断的地方保留监督。通过支持跨多种输入类型的可定制提取工作流,该软件为材料研究以及可能多个学科的可扩展、可靠的数据整理提供了一个灵活的框架。用户有责任确保所有输入、提取工作流和下游使用符合适用的出版商协议、版权和许可条款、机构政策以及数据使用要求。

英文摘要

SALSA (Semi-Autonomous Literature Summarization Assistant) is an open- source, human-in-the-loop platform for extracting structured scientific datasets from multimodal literature sources. The software combines document parsing, large language models, optical character recognition, computer vision, figure digitization, and user-guided correction tools to recover structured information from text, tables, figures, and captions. Users can configure extraction stages, define dataset schemas, perform interventions on digitized figures, and export verified data for downstream analysis and machine learning. SALSA is designed to automate repetitive literature curation tasks while preserving oversight where expert judgment is required. By supporting customizable extraction workflows across diverse input types, the software provides a flexible framework for scalable, reliable data curation for materials research, and potentially across several disciplines. Users are responsible for ensuring that all inputs, extraction workflows, and downstream uses comply with applicable publisher agreements, copyright and licensing terms, institutional policies, and data-use requirements.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑