边缘开放,中心捕获:llama.cpp 与本地 AI 推理的政治经济学
Open at the Edge, Captured at the Center: llama.cpp and the Political Economy of Local AI Inference
AI总结:
该研究以 llama.cpp 为对象,分析本地 AI 推理的政治经济动态,发现本地推理拓宽参与度却将控制权向基础设施相关方转移,呼吁政策关注推理基础设施以维持云外 AI 开放性。
AI中文摘要:
开放人工智能学术研究聚焦于模型发布与云生态系统,却基本未考察让开放权重模型可在用户自有设备上运行的本地推理基础设施。我们通过对 llama.cpp 的混合方法分析填补这一空白,结合 2023 年 3 月至 2026 年 3 月的 7681 次合并拉取请求、仓库讨论、企业声明及贡献者博客开展研究。我们表明,本地推理拓宽了执行环节的参与度,却将控制权转移至使执行成为可能的基础设施。通过硬件后端、模型集成工作,以及 Hugging Face 在 2026 年 2 月对该项目的收购,我们记录了控制权如何转移至硬件供应商、模型分发商及核心维护者,而模型所有者与个体贡献者则承担着让模型可运行的成本。这些动态表明,要在云之外保持开放性,需关注使模型可运行的基础设施,而非仅关注模型本身。这要求政策机制——对格式依赖与供应商影响力的分析、模型兼容性要求,以及对推理工具的持续公共资助——这些机制需超出模型发布条件,延伸至基础设施层面。
英文摘要:
Critical scholarship on open AI has focused on model releases and cloud ecosystems, leaving the local inference infrastructure that makes open-weight models runnable on user-owned devices largely unexamined. We address this gap through a mixed-methods analysis of llama$.$cpp, combining 7,681 merged pull requests from March 2023 through March 2026 with repository discussions, corporate statements, and contributor blogs. We show that local inference broadens participation at execution while relocating capture into the infrastructure that makes execution possible. Through hardware backends, model integration labor, and Hugging Face's February 2026 absorption of the project, we document how control shifts to hardware vendors, model distributors, and core maintainers while model owners and individual contributors bear the cost of making models runnable. These dynamics suggest that preserving openness outside the cloud requires attention to the infrastructure that makes models runnable, not just to the models themselves. This calls for policy mechanisms--analysis of format dependencies and vendor influence, model compatibility requirements, and sustained public funding for inference tooling--that extend beyond model release conditions to the infrastructure layer.