arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24807cs.SEcs.AI

基于大语言模型的自动模型卡片生成

Automatic Model Card Generation Using an LLM

  • University of Alberta(阿尔伯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

Tajkia Rahman Toma, Balreet Grewal, Cor-Paul Bezemer

AI总结:

本研究提出基于LLM的MCTidy与MCGenie,前者重组模型卡片为标准化模板,后者从仓库数据直接生成,经48个Hugging Face模型验证,可实现标准化、可扩展的模型卡片文档生成。

AI中文摘要:

模型卡片是汇总机器学习模型关键信息的结构化文档,旨在提升透明度、可用性与可问责性。然而,它们往往缺乏统一结构,且许多模型未提供模型卡片,导致比较与解释困难。本文贡献有二:其一,提出MCTidy,一种基于大语言模型(LLM)的方法,可将现有模型卡片重组为标准化模板,以提升清晰度与可比性;其二,引入MCGenie,一种基于LLM的系统,可直接从模型仓库数据生成模型卡片。我们将MCTidy应用于48张Hugging Face模型卡片,评估信息保留度、章节对齐度、幻觉情况与稳定性,结果显示其信息保留度高、文本损失极小、章节分配准确、幻觉仅出现在描述性章节且频率低,多次运行稳定性强。我们通过为相同48个模型生成模型卡片,评估MCGenie的语义相似度、事实正确性与对输入资源的敏感性,生成的模型卡片语义相似度高(均值约0.9),超半数完全正确,其余多数错误为 minor(次要),生成质量高度依赖支持资源(尤其是关联论文)的可用性。总体而言,本研究证明了基于LLM的方法实现可扩展、标准化模型卡片文档的潜力。

英文摘要:

Model cards are structured documents that summarize key information about machine learning models to improve transparency, usability, and accountability. However, they often lack a consistent structure, and many models provide no model cards, making comparison and interpretation difficult. This paper presents two contributions. First, we propose MCTidy, an LLM-based approach that reorganizes existing model cards into a standardized template to improve clarity and comparability. Second, we introduce MCGenie, an LLM-based system that generates model cards directly from model repository data. We apply MCTidy to 48 Hugging Face model cards and evaluate information retention, section alignment, hallucination, and stability. Our findings show high information retention with minimal textual loss, accurate section assignment, rare hallucinations primarily in descriptive sections, and strong stability across runs. We assess MCGenie by generating model cards for the same 48 models and assessing semantic similarity, factual correctness, and sensitivity to input resources. The generated model cards achieved high semantic similarity (mean around 0.9); over half were fully correct, and most remaining errors were minor. Generation quality depended strongly on the availability of supporting resources, particularly associated papers. Overall, our findings demonstrate the potential of LLM-based methods to enable scalable, standardized model card documentation.

↑