arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

INRIA 数据湖:用于 HAL 应用于软件提及跟踪的通用且可扩展的管道生态系统

The INRIA DataLake: A Generic and Scalable Ecosystem of Pipelines for HAL Applied to Software Mentions Tracking

Luca Foppiano, Vipul Gupta, Samuel Scalbert, Estelle Nivault, Kumar Guha, Yannick Barborini, Alain Monteil, Laurent Romary

arXiv 2607.09824首次发表:更新:

AI 中文总结

该论文介绍 INRIA 数据湖项目,它提供可扩展且互联的管道生态系统。利用 Grid'5000/ABACA 基础设施,通过提取科学文章中软件提及并可视化的用例展示。结果表明系统能高效处理文献,支持用户验证与外部互操作,为开放科学做贡献。

AI 中文摘要

研究存储库包含大量科学知识,但获取结构化文章和诸如数据集或软件元数据等专业信息仍然有限。本文介绍了 INRIA 数据湖项目,它提供了一个可扩展且相互连接的管道生态系统,用于准备科学文献、提取结构化信息并应用专业处理。利用大规模共享基础设施 Grid'5000/ABACA,通过一个具体用例展示了该生态系统:从每日存入的科学文章中提取软件提及,并在 HAL 研究门户中验证后进行可视化。结果表明该系统能高效处理大量科学文献,同时支持用户验证以及与外部系统的互操作性。该项目旨在通过整合额外管道并跨研究团队共享准备工作来发展,已通过提高研究软件的可见性和跟踪为开放科学做出贡献。

英文摘要

Research repositories contain a large amount of scientific knowledge, but access to structured articles and specialised information, such as datasets or software metadata, remains limited. In this paper, we present the INRIA DataLake project, which provides an ecosystem of scalable and interconnected pipelines for preparing scientific literature, extracting structured information, and applying specialised treatments. Using a large-scale shared infrastructure, Grid'5000/ABACA, we demonstrate our ecosystem through a concrete use case: extracting software mentions from scientific articles deposited daily and visualising them after validation in the HAL research portal. Our results show that the system can efficiently process large volumes of scientific literature while supporting user validation and interoperability with external systems. Designed to grow by integrating additional pipelines and sharing the preparation effort across research groups, this project already contributes to open science through improved visibility and tracking of research software.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑