无审查开放权重模型:以重分发作为持久层
Uncensored Open-weight Models: Redistribution as the Persistence Layer
查看机构详情
- a Labs(10a实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究分析了移除安全护栏的无审查开放权重模型生态,识别了模型数量、重分发情况及相关应用,发现部分集成模型的应用存在恶意风险。
中文摘要 AI 辅助
一个不断扩张的行为主体生态系统正在移除开放权重AI模型内置的安全护栏。我们通过识别关键生产者、下游复现案例及新兴应用,对该生态系统进行了分析。2024年1月至2026年3月间,我们在HuggingFace平台上识别出3471个原始无审查模型,每个模型平均被重新打包2.4次;3个主体占全部8164次压缩重分发的52%。这些模型经量化后,会被镜像到不同账户、格式及Ollama等注册中心,从而在上游移除后仍能保留,且更便于下游部署。在识别出的1643个集成无审查大语言模型(ULLMs)的GitHub应用中,25%被归类为明确恶意应用。
英文摘要
A rapidly expanding ecosystem of actors is removing built-in safety guardrails from open-weight AI models. We profile this ecosystem by identifying key producers, downstream reproductions, and emerging applications. Between January 2024 and March 2026, we identified 3,471 original uncensored models on HuggingFace, each repackaged an average of 2.4 times; three actors account for 52% of all 8,164 compressed redistributions. Once quantized and mirrored across separate accounts, formats, and registries such as Ollama, these models persist regardless of upstream removal and become easier to deploy downstream. Of the 1,643 identified GitHub applications integrating uncensored large language models (ULLMs), 25% were classified as explicitly malicious.