arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向LLM教师蒸馏标注的可扩展流水线:工作窃取作业调度与内存感知GPU并发控制

A Scalable Pipeline for LLM-Teacher Distillation Labeling: Work-Stealing Job Scheduling and Memory-Aware GPU Concurrency

Ravi Satya Durga Prasad Yenugula

arXiv 2608.15975首次发表:更新:

AI 中文总结

该研究提出了一种用于LLM教师蒸馏标注的可扩展流水线,通过工作窃取作业调度、内存感知GPU并发控制及重标注基准方法,解决了标注的质量成本与GPU集群负载问题,性能优于静态分片且可复现。

AI 中文摘要

用大型语言模型(LLM)教师对大型文本语料库进行标注,已成为规模化获取训练数据的可行途径。当数据量达到数百万项时,对每一批次进行人工标注不可行,核心问题在于:教师每单位成本能获得的标注质量,以及如何在负载倾斜、易出现故障的情况下维持GPU工作集群的繁忙状态。本文提出了一种简单、可复现的流水线,同时解决这两个问题。其一,工作窃取环形池:每个工作节点拥有一个队列,先处理自身队列,再从环形后继节点窃取任务;通过原子条件写入实现任务的一次仅声明机制,通过陈旧声明清理实现容错。该声明协议仅需存储层提供比较并设置(compare-and-set)原语,我们在单个SQLite文件上实现该机制,使参考实现无依赖,且可在单台机器上复现实验。其二,内存感知并发规则:根据GPU能容纳的模型副本数量调整节点内并行度,使同一代码可在不同设备规格上安全运行。其三,重标注基准测试方法:教师对已包含黄金标签的公开数据集进行重标注,将质量评估转化为一致性测量,成本则由测得的吞吐量得出。在倾斜负载下,该环形池的吞吐量是静态分片的3.4倍,零倾斜时与静态分片持平;在运行中途终止一半工作节点时,环形池在2000个任务中损失0个,而静态分片损失953个;该方法还得出了指令调优教师在反讽和情感分析任务上的质量与成本对应关系。所有实验均在公开数据和商用硬件上运行,代码、测试及运行日志已公开。

英文摘要

Labeling large text corpora with LLM teachers has become a practical route to training data at scale. At millions of items, hand-labeling every batch is not feasible, and two questions dominate: what label quality a teacher buys per dollar, and how to keep a fleet of GPU workers busy under skewed, failure-prone workloads. We present a simple, reproducible pipeline that addresses both. First, a work-stealing ring pool: each worker owns a queue, drains it first, and then steals from ring successors, with exactly-once task claims via atomic conditional writes and crash tolerance via stale-claim sweeping. The claim protocol requires only a compare-and-set primitive from its storage layer; we implement it on a single SQLite file, which makes the reference implementation dependency-free and the experiments reproducible on one machine. Second, a memory-aware concurrency rule that sizes per-node parallelism by how many model copies fit on the GPU, so the same code runs safely across device sizes. Third, a relabeling benchmark methodology in which the teacher relabels a public dataset that already has gold labels, so quality reduces to an agreement measurement and cost follows from measured throughput. Under skewed load the pool sustains up to 3.4 times the throughput of static sharding while matching it at zero skew, loses 0 of 2,000 tasks when half the workers are killed mid-run (static sharding loses 953), and yields measured quality and cost points for an instruction-tuned teacher on irony and sentiment tasks. All experiments run on public data and commodity hardware; code, tests, and run logs are released.

Comments8 pages, 1 figure, 3 tables. Code, tests, and all run artifacts: https://github.com/rsdpyenugula/hybrid-labeling-training

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑