arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可见但尚未可整理:表征紧凑型与衍生型开源大语言模型(LLM)制品的可整理性

Visible but Not Yet Curatable: Characterizing the Curatability of Compact and Derived Open LLM Artifacts

Yiyi Lu, Yilai Qian, Yucheng Jin

arXiv 2608.28819首次发表:更新:

发表机构

Duke Kunshan University(杜克昆山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出框架表征开源LLM紧凑型与衍生型制品的可整理性,发现仅6.1%的相关记录具备高可整理性,并为模型中心等机构提出改进建议。

AI 中文摘要

开源大语言模型(LLM)研究日益产出紧凑型与衍生型制品,如适配器(adapters)、量化检查点(quantized checkpoints)、合并模型(merged models)及蒸馏变体(distilled variants),这些制品分布于论文、模型中心(model hubs)、模型卡片(model cards)、代码仓库及发布声明中。尽管这些制品公开可见,但数字图书馆往往缺乏足够证据将其作为连贯的学术对象进行识别、保存与引用。我们提出一个框架,将可整理性概念化为分布式学术记录的记录级属性,并通过四个证据维度将其可操作化:制品身份、学术关联、上游证据及发布资产。在该框架指导下,我们利用2026年5月的191375个公开Hugging Face仓库快照及2214篇学术论文核心语料,开展了首个针对开源LLM可整理性的全规模表征研究。结果显示存在显著的“可见-可整理”漏斗效应:90.7%的论文记录包含至少一项有用的整理信号,但仅18.1%的记录兼具可用的上游证据与具体的发布证据,仅有6.1%的记录提供了足够协调的证据以支持高可整理性记录。基于这些发现,我们推导了一个包含七个字段的最小可整理记录,并明确了模型中心、学术索引及数字图书馆的互补责任,为改进开源LLM制品的保存与书目控制提供了实践指导。

英文摘要

Open Large Language Model (LLM) research increasingly produces compact and derived artifacts, such as adapters, quantized checkpoints, merged models, and distilled variants, that are distributed across papers, model hubs, model cards, code repositories, and release statements. Although these artifacts are publicly visible, digital libraries often lack sufficient evidence to identify, preserve, and cite them as coherent scholarly objects. We introduce a framework that conceptualizes curatability as a record-level property of distributed scholarly records and operationalizes it through four evidence dimensions: artifact identity, scholarly linkage, upstream evidence, and release assets. Guided by this framework, we conduct the first collection-scale characterization of open LLM curatability using a May 2026 snapshot of 191,375 public Hugging Face repositories and a core corpus of 2,214 scholarly papers. Our results reveal a pronounced visibility-to-curatability funnel. While 90.7% of paper records contain at least one useful curation signal, only 18.1% combine usable upstream evidence with concrete release evidence, and only 6.1% provide sufficiently coordinated evidence to support high-curatability records. Based on these findings, we derive a minimal seven-field curatable record and complementary responsibilities for model hubs, scholarly indexes, and digital libraries, providing practical guidance for improving the preservation and bibliographic control of open LLM artifacts.

Comments11 pages, 5 figures. Accepted at the ACM/IEEE Joint Conference on Digital Libraries (JCDL 2026). Yiyi Lu and Yilai Qian contributed equally; Yucheng Jin is the corresponding author. Code and results: https://github.com/AndyLu666/VNYC-Open_Source_LLM_Study

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑