发表机构
Xpeng Inc.(小鹏汽车)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
IronLLM提出紧凑边缘语言模型,通过混合注意力、轻量级多令牌预测和策略蒸馏,在设备端实现高效推理,性能媲美更大模型。
AI 中文摘要
我们提出了IronLLM-0.6B,一个具有654M参数的语言模型,专为高效的设备端推理而设计。IronLLM-0.6B结合了混合注意力架构与X-MTP,一种轻量级的共享KV多令牌预测设计,该设计消除了逐层KV缓存重放,并采用轻量级验证头实现无回滚草稿,实现了1.48倍的解码加速。该模型在约6.2万亿个令牌上使用面向质量的数据管道进行预训练,并进一步通过多领域策略蒸馏进行后训练,以整合来自领域专家教师的能力。为了更好地满足设备端场景的低延迟要求,IronLLM-0.6B采用了仅指令设计。评估表明,IronLLM-0.6B在性能上与更大的模型如Qwen3.5-0.8B和MiniCPM5-1B相比具有竞争力,同时在许多任务上产生更简洁的响应。我们进一步提出了IronLLM-0.6B-Light,它用动态Tanh替换了RMSNorm,并简化了若干计算昂贵的组件,以提高推理和量化效率。总之,IronLLM模型为资源受限的部署提供了有效的性能-效率权衡。
英文摘要
We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-free drafting, achieving a 1.48x decoding speedup. The model is pretrained on approximately 6.2 trillion tokens using a quality-oriented data pipeline and is further post-trained with Multi-Domain On-Policy Distillation to integrate capabilities from domain-specialized teachers. To better meet the low-latency requirements of on-device scenarios, IronLLM-0.6B adopts an Instruct-Only design. Evaluations show that IronLLM-0.6B achieves competitive performance relative to larger models such as Qwen3.5-0.8B and MiniCPM5-1B, while producing more concise responses on many tasks. We further present IronLLM-0.6B-Light, which replaces RMSNorm with Dynamic Tanh and simplifies several computationally expensive components to improve inference and quantization efficiency. Together, the IronLLM models provide an effective performance-efficiency trade-off for resource-constrained deployment.
CommentsTechnical report