对Hugging Face模型中人工智能物料清单完整性的大规模测量
A Large-Scale Measurement of AI Bill of Materials Completeness in Hugging Face Models
浏览论文内容
中文总结 AI 辅助
本文以Hugging Face模型库为案例,研究人工智能物料清单(AIBOMs)完整性。通过检查大量工件评估AIBOMs在多方面表现,发现其结构覆盖完整,但特定文档完整性有限,部分字段缺失。研究结果推动改进相关实践以促进更完整AIBOMs的生成和采用。
中文摘要 AI 辅助
预训练机器学习模型助力开发者构建系统,但模型库关于模型出处、许可等的机器可读文档常不完整,造成人工智能供应链透明度和治理差距。人工智能物料清单(AIBOMs)可解决此问题。本文以公共Hugging Face模型库为案例研究AIBOM完整性,即库为机器可读人工智能供应链文档提供AIBOM相关信息的程度。通过检查约97.5K个AIBOM工件评估生成的AIBOMs在多方面的情况,结果显示其结构覆盖完整,但特定于人工智能的文档完整性有限,部分字段表示薄弱或缺失。研究结果推动改进模型卡实践等以促进更完整AIBOMs的生成和采用。
英文摘要
Pretrained machine learning (ML) models help developers build ML-intensive software systems without training models from scratch. However, model repositories often provide incomplete machine-readable documentation about model provenance, licenses, datasets, limitations, and external references, creating transparency and governance gaps across the AI supply chain. Artificial Intelligence Bills of Materials (AIBOMs) address these gaps by documenting AI artifacts, including models, metadata, licenses, datasets, model-card information, and external references. Taking public Hugging Face (HF) model repositories as a case study, this paper empirically investigates AIBOM completeness, defined as the extent to which repositories provide AIBOM-relevant information for machine-readable AI supply-chain documentation. We examine approximately 97.5K AIBOM artifacts to assess the extent to which generated AIBOMs: (i) contain required structural and metadata fields, (ii) represent model identity, license, and external-reference information, (iii) capture model-card documentation such as datasets, limitations, safety-risk assessment, and environmental information, and (iv) vary in documentation coverage across repository and artifact characteristics such as task, license availability, dataset declaration, model family, and paper reference. Results indicate that generated AIBOMs provide complete coverage of required AIBOM structure but limited AI-specific documentation completeness. Required fields are fully represented, but model-card, metadata, responsible-use, environmental, limitation, and meaningful-description fields remain weakly represented or missing across generated artifacts. Our findings motivate improved model-card practices, repository-level traceability, and automated AIBOM validation to advance the generation and adoption of more complete AIBOMs.
发表机构
- Department of Computer Science, The University of Alabama(计算机科学系,阿拉巴马大学)
机构由 AI 辅助整理,请以论文原文为准。