arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02211cs.CEcond-mat.mtrl-sci

datascribe_api:借助DataScribe实现数据驱动的材料发现

datascribe_api: Enabling Data-Driven Materials Discovery with DataScribe

Doğuhan Sarıtürk, Raymundo Arróyave, Vahid Attari

首次发表
浏览论文内容

中文总结 AI 辅助

datascribe_api是一个Python库和CLI,提供对DataScribe Cloud的编程访问,整合用户数据表和精选材料数据库,通过类型化接口和九个查询端点简化数据检索,支持pandas转换和FAIR溯源,促进数据驱动的材料发现。

中文摘要 AI 辅助

datascribe_api是一个Python库和命令行界面(CLI),为研究人员提供对DataScribe Cloud的程序化访问,DataScribe Cloud是一个用于材料设计与发现的AI原生平台。DataScribe整合了两类数据服务:用户管理的科学数据表,研究团队在其中存储实验测量、模拟输出和衍生特征;以及来自Materials Project、AFLOW和OQMD的精选材料科学数据库。datascribe_api通过一个统一的、类型化的Python接口暴露这两类服务,使研究人员无需编写自定义HTTP代码即可检索、筛选和分析数据。该库自动处理身份验证、连接管理和瞬时网络错误。它将九个查询端点暴露为普通的Python方法调用。筛选条件使用标准的Python比较语法,并以列表形式传递,使查询逻辑可读且无需特定格式的样板代码。Pydantic模型在解析时验证所有API响应,因此模式不匹配会以异常形式显现,而不是静默的错误值。每个结果集合通过单次调用转换为pandas DataFrame,使检索到的数据可直接用于机器学习和统计分析。每个材料搜索结果都带有标识其来源数据库的溯源字段,满足FAIR可追溯性要求。该库包含一个CLI,将全部九个端点暴露为子命令,使用Typer进行参数解析,使用Rich进行终端输出。每个子命令都支持机器可读的JSON输出,用于shell管道和LLM代理。为Python 3.11及更高版本预构建的wheel包已在PyPI上以GPLv3许可证发布,支持Linux、macOS和Windows。

英文摘要

datascribe_api is a Python library and command-line interface (CLI) that gives researchers programmatic access to DataScribe Cloud, an AI-native platform for materials design and discovery. DataScribe integrates two data services: user-managed scientific data tables, where research groups store experimental measurements, simulation outputs, and derived features; and curated materials science databases from the Materials Project, AFLOW, and OQMD. datascribe_api exposes both through a single, typed Python interface, so researchers can retrieve, filter, and analyze data without writing custom HTTP code. The library handles authentication, connection management, and transient network errors automatically. It exposes nine query endpoints as ordinary Python method calls. Filter conditions use standard Python comparison syntax and are passed as a list, keeping query logic readable without format-specific boilerplate. Pydantic models validate all API responses at parse time, so schema mismatches surface as exceptions rather than silent wrong values. Each result collection converts to a pandas DataFrame in a single call, so retrieved data is directly available for machine learning and statistical analysis. Every materials search result carries a provenance field identifying its source database, satisfying FAIR traceability requirements. The library includes a CLI that exposes all nine endpoints as subcommands, using Typer for argument parsing and Rich for terminal output. Every subcommand supports machine-readable JSON output for shell pipelines and LLM agents. Pre-built wheels for Python 3.11 and later are published on PyPI under the GPLv3 license for Linux, macOS, and Windows.

发表机构

  • Department of Materials Science
  • Engineering, Texas A\&M University, USA
  • J. Mike Walker '66 Department of Mechanical Engineering, Texas A\&M University, USA
  • Wm Michael Barnes '64 Department of Industrial
  • Systems Engineering, Texas A\&M University, USA

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑