到达即损坏:公共模型注册表中静默缺陷的LLM工件及其捕获方法
Broken on Arrival: Silently Defective LLM Artifacts in Public Model Registries and How to Catch Them
浏览论文内容
中文总结 AI 辅助
本研究通过大规模测试发现公共模型注册表中存在静默缺陷的LLM工件,提出quantcheck验收测试工具,并主张模型注册表应引入类似包注册表的验收门槛。
中文摘要 AI 辅助
开发者越来越多地通过从公共注册表拉取量化的GGUF工件来本地运行大型语言模型,然而在分发管道中,没有任何环节会在这些转换到达用户之前对其进行功能测试。我们执行了327个量化且具备代码能力的模型工件:其中305个来自官方Ollama库,涵盖15个模型系列,覆盖8GB及以下所有合格量化级别;另外22个来自HuggingFace上下载量最高的社区仓库。每个工件都运行了一个15任务冒烟测试套件,该套件经过校准,使得健康工件能够通过而已知损坏的工件会失败;可疑工件随后接受完整的164任务评估、第二个推理后端、独立分发者对同一模型和量化级别的转换作为仲裁,以及对于社区文件,在其自身模板下重新测试。官方库中存在五个静默缺陷工件,包括一批四个Qwen2.5-Coder-3B转换和一个phi3.5-mini转换,这些工件在两个后端上均无法解决164个任务中的任何一个,也无法通过冒烟测试套件中的任何一项,而同一模型的独立转换却能正常工作:占官方工件的1.6%,占29个模型与规模转换组中的2个。仲裁链为小型模型工件洗清了罪名,这些工件若被简单阈值判定会因极端量化导致崩溃而被误判为损坏;同时,仲裁链还暴露了两个较旧的社区转换,它们在CUDA上严重退化但在Metal上通过:这不是文件缺陷而是后端依赖的失败,这是目前没有任何注册表测试的第三种现象。两个已确认的缺陷产生的输出,其表面统计量处于健康范围内,任何低噪声启发式方法若不执行测试都无法察觉。我们发布了审计数据集、quantcheck验收测试工具以及每个已确认缺陷的披露报告(见https链接),并主张模型注册表需要像包注册表已经运行的那样设置验收门槛。
英文摘要
Developers increasingly run large language models locally by pulling quantized GGUF artifacts from public registries, yet nothing in the distribution pipeline functionally tests these conversions before they reach users. We executed 327 quantized code-capable model artifacts: 305 from the official Ollama library, spanning 15 model lines at every eligible quantization level at or under 8 GB, and 22 from the most-downloaded community repositories on HuggingFace. Each ran a 15-task smoke suite calibrated so that healthy artifacts pass while a known-broken one fails; suspects then faced full 164-task evaluation, a second inference backend, an independent distributor's conversion of the same model and quantization as referee, and, for community files, re-testing under the artifact's own template. The official library carries five silently defective artifacts, a batch of four Qwen2.5-Coder-3B conversions and one phi3.5-mini conversion, that solve zero of 164 tasks and zero of the smoke suite on both backends while independent conversions of the same models work: 1.6% of official artifacts, 2 of 29 model-and-size conversion groups. The adjudication chain cleared small-model artifacts that a naive threshold would condemn as broken when they are merely collapsed by extreme quantization, and it exposed two older community conversions that degrade badly on CUDA yet pass on Metal: not defective files but backend-dependent failures, a third phenomenon no registry currently tests for. Two confirmed defects produce output whose surface statistics sit inside the healthy range, invisible to any low-noise heuristic short of execution. We release the audit dataset, the quantcheck acceptance-testing tool, and disclosure reports for every confirmed defect (https://github.com/aditi-p31/quantcheck), and argue that model registries need the acceptance gate that package registries already run.
发表机构
- Independent Researcher Email
机构由 AI 辅助整理,请以论文原文为准。