arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07226cs.ARcs.LG

基于Tailscale的双节点NVIDIA DGX Spark:用于分布式大语言模型训练和网络威胁情报微调的远程访问测试平台

Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-Tuning

Vasanth Iyer

首次发表
浏览论文内容

中文总结 AI 辅助

本研究展示了双节点NVIDIA DGX Spark通过Tailscale VPN实现分布式NanoChat预训练,构建网络安全微调数据集并验证其效果,同时该集群可支持研究与教学,证明适度本地基础设施的可行性。

中文摘要 AI 辅助

小型AI系统让本地语言模型实验变得越来越容易,但桌面级加速器上的多节点训练的实际证据仍然有限。本报告展示了在两台NVIDIA DGX Spark系统上进行分布式NanoChat预训练的概念验证部署,每台系统配备GB10 Grace Blackwell片上系统和128 GB统一内存,通过Tailscale网状VPN进行远程管理,并通过专用200 Gb/s QSFP56直连光纤链路连接以进行训练。PyTorch的torchrun、DDP和NCCL被配置为每个节点一个进程,采用深度为20的NanoChat模型,每个节点的本地批量大小为32,上下文长度为2048个token,从而每步的全局批量为131072个token。该运行维持了约69.4秒的步时间(约1890个token/秒),在四天内处理了约6.53亿个token。我们记录了链路配置、容器设置、接口绑定、触发NCCL超时的步零评估错误、 checkpointing以及故障排除经验,作为小型实验室的可重复性参考。我们还从77份CISA公告中构建了一个网络安全微调数据集(338个训练对话、37个验证对话),并进行了17个问题的保留评估,使用Ollama托管的大语言模型评判器比较基线SFT checkpoint与CTI增强型checkpoint。CTI特定类别表现提升,而通用知识类别表现下降,整体从0-10分制的2.06小幅变化至2.29。同一集群还支持400级AI课程(CS 426)和CBS 255中用于CompTIA Security+ POGIL活动的查询引擎,表明适度的本地基础设施可同时支持研究和教学。本研究确立了可行性而非缩放效率主张,因为用于比较的单节点吞吐量是估计值,而非在匹配条件下测量的。操作手册和脚本可用(见代码可用性)。

英文摘要

Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class accelerators remains limited. This report presents a proof-of-concept deployment of distributed NanoChat pretraining across two NVIDIA DGX Spark systems, each with a GB10 Grace Blackwell system-on-chip and 128 GB of unified memory, administered remotely over a Tailscale mesh VPN and connected for training by a dedicated 200 Gb/s QSFP56 direct fiber link. PyTorch torchrun, DDP, and NCCL were configured with one process per node, a depth-20 NanoChat model, a local batch size of 32 per node, and a 2,048-token context, giving a global batch of 131,072 tokens per step. The run sustained a step time of about 69.4 s (about 1,890 tokens/s), processing about 653 million tokens over four days. We document link configuration, container setup, interface binding, a step-zero evaluation bug that triggered NCCL timeouts, checkpointing, and troubleshooting lessons, as a reproducibility reference for small labs. We also built a cybersecurity fine-tuning dataset from 77 CISA advisories (338 training, 37 validation conversations) and ran a 17-question held-out evaluation comparing a baseline SFT checkpoint against a CTI-augmented checkpoint with an Ollama-hosted LLM judge. CTI-specific categories improved while general-knowledge categories regressed, for a small overall change from 2.06 to 2.29 on a 0-10 scale. The same cluster supports a 400-level AI course (CS 426) and a query engine for CompTIA Security+ POGIL activities in CBS 255, showing modest local infrastructure can serve both research and teaching. The study establishes feasibility rather than a scaling-efficiency claim, since single-node throughput used for comparison was estimated, not measured under matched conditions. Runbook and scripts are available (see Code Availability).

发表机构

  • Grambling State University(格兰布林州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑