arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Lightning Weave:通过能力组合提升推理模型的精度-效率前沿

Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition

Yecheng Wu, Song Han, Han Cai

arXiv 2609.14708首次发表:更新:

发表机构

Massachusetts Institute of Technology; NVIDIA(麻省理工学院; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Lightning Weave通过在线策略蒸馏组合独立后训练模型的精度与效率能力,在数学和代码基准上显著提升精度并减少响应词元,实现最先进的精度-效率前沿。

AI 中文摘要

高效推理的一个核心目标是改善精度-效率前沿。然而,同时提高推理精度和推理效率可能具有挑战性,因为这两个目标可能偏好不同的推理行为。独立后训练的模型已经在精度和效率方面展现出各自的优势。我们提出了Lightning Weave,一种后训练框架,通过在线策略蒸馏将独立学习到的能力提取并组合到单个学生模型中。每个获得的能力由模型后训练前后的策略偏移表示,即从后训练前的模型到最终专家模型的策略变化。Lightning Weave在共享的学生词元状态上组合对齐的对数比率偏移,并使用Tilted-Target DOPD将缓存信号转换为稳定的学习目标。每个锚点对仅对缓存轨迹评分一次,使得后续学生训练无需同时运行多个实时锚点模型。在数学和代码领域的多种学生模型和基准测试中,Lightning Weave显著优于基础学生模型,并达到了最先进的精度-效率前沿。在Qwen3.5-4B上,它将HMMT 2025的准确率从59.2%提升至64.0%,同时响应词元减少10.7%;将LiveCodeBench v5的准确率从41.7%提升至54.2%,同时响应词元减少9.6%。调整锚点信号的相对强度可产生强经验性的精度-效率帕累托前沿。这些结果确立了Lightning Weave作为通过能力组合实现高效推理的新实用路径。代码即将发布。

英文摘要

A core goal of efficient reasoning is to improve the accuracy-efficiency frontier. However, jointly improving reasoning accuracy and inference efficiency can be challenging, as the two objectives can favor different reasoning behaviors. Independently post-trained models already offer distinct strengths in accuracy and efficiency. We introduce Lightning Weave, a post-training framework that extracts and composes these independently learned capabilities in a single student through on-policy distillation. Each acquired capability is represented by the policy shift from the model before post-training to the resulting specialist. Lightning Weave combines aligned log-ratio shifts at shared student token states and uses Tilted-Target DOPD to convert the cached signals into a stable learning target. Each anchor pair scores the cached trajectories once, enabling subsequent student training without serving multiple live anchor models concurrently. Across diverse student models and benchmarks in mathematics and code, Lightning Weave substantially improves upon the base students and achieves a state-of-the-art accuracy-efficiency frontier. On Qwen3.5-4B, it raises HMMT 2025 accuracy from 59.2% to 64.0% with 10.7% fewer response tokens, and LiveCodeBench v5 accuracy from 41.7% to 54.2% with 9.6% fewer response tokens. Adjusting the relative strengths of the anchor signals yields a strong empirical accuracy-efficiency Pareto frontier. These results establish Lightning Weave as a new practical route to efficient reasoning through capability composition. Code is released at https://github.com/jet-ai-projects/Lightning-Weave.

CommentsWork in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑