AI 中文总结
本文介绍使用Nix构建混合HPC/AI软件栈的经验,利用Nix特性解决传统方法的依赖等问题,统一多语言管理并生成部署容器,同时讨论相关权衡与Nixpkgs的ML包缺口。
AI 中文摘要
在生产型超级计算机的约束下,HPC的可复现性仍存在困难:无root权限、网络受限,且软件栈日益涵盖C/C++、Fortran、Python、MPI及GPU运行时。基于环境模块和Conda的传统方法需手动定位依赖,会将系统库泄露到构建中,且无法跨项目组合。容器有助于部署,但本身不保证可复现性。本文报告使用Nix构建混合HPC/AI软件栈的经验,涵盖无root权限的工作站本地开发,以及生产集群上作为Apptainer镜像的远程部署。Nix一致的包布局、完整的环境隔离和基于flake的组合,解决了遇到的依赖发现、泄露和组合问题,同时在单个声明式规范下统一了C/C++和Python的管理,该规范还能生成部署容器。本文讨论了与Spack和Guix的权衡,通过CMake预设解决的开发与生产分离问题,以及Nixpkgs中ML包覆盖的当前缺口。
英文摘要
Reproducibility in HPC remains difficult under the constraints of production supercomputers: no root access, limited internet, and software stacks that increasingly span C/C++, Fortran, Python, MPI, and GPU runtimes. Traditional approaches based on environment modules and Conda require manual intervention to locate dependencies, leak system libraries into builds, and fail to compose across projects. Containers help with deployment but do not by themselves guarantee reproducibility. We report on our experience building a hybrid HPC/AI software stack with Nix, covering local development on a workstation without root and remote deployment as an Apptainer image on a production cluster. Nix's consistent package layout, full environment isolation, and flake-based composition resolve the dependency discovery, leakage, and composition problems we encountered, while unifying C/C++ and Python management under a single declarative specification that also generates the deployment container. We discuss trade-offs against Spack and Guix, the development-versus-production split addressed via CMake presets, and current gaps in ML package coverage in Nixpkgs.
Journal ref1st Workshop in Sustainable Practices for Reproducibility in HPC, Jun 2026, Hamburg, Germany