arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IronLLM:锻造面向实时具身智能的紧凑边缘原生语言模型

IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence

Changdi Yang, Fengquan Jiao, Haochih Lin, Haoran Yang, Jing Xiao, Liangyu Huo, Suxin Lu, Tiance Chen, Wei Liu, Yinggan Xu, Yunxiang Lu, Zai Zheng, Zhirui Xie, Zhongyang Che, Ziyan Tang, Zuoxiang Zhao, Jian Yao

arXiv 2609.36860首次发表:更新:

发表机构

Xpeng Inc.(小鹏汽车)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

IronLLM提出紧凑边缘语言模型,通过混合注意力、轻量级多令牌预测和策略蒸馏,在设备端实现高效推理,性能媲美更大模型。

AI 中文摘要

我们提出了IronLLM-0.6B,一个具有654M参数的语言模型,专为高效的设备端推理而设计。IronLLM-0.6B结合了混合注意力架构与X-MTP,一种轻量级的共享KV多令牌预测设计,该设计消除了逐层KV缓存重放,并采用轻量级验证头实现无回滚草稿,实现了1.48倍的解码加速。该模型在约6.2万亿个令牌上使用面向质量的数据管道进行预训练,并进一步通过多领域策略蒸馏进行后训练,以整合来自领域专家教师的能力。为了更好地满足设备端场景的低延迟要求,IronLLM-0.6B采用了仅指令设计。评估表明,IronLLM-0.6B在性能上与更大的模型如Qwen3.5-0.8B和MiniCPM5-1B相比具有竞争力,同时在许多任务上产生更简洁的响应。我们进一步提出了IronLLM-0.6B-Light,它用动态Tanh替换了RMSNorm,并简化了若干计算昂贵的组件,以提高推理和量化效率。总之,IronLLM模型为资源受限的部署提供了有效的性能-效率权衡。

英文摘要

We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-free drafting, achieving a 1.48x decoding speedup. The model is pretrained on approximately 6.2 trillion tokens using a quality-oriented data pipeline and is further post-trained with Multi-Domain On-Policy Distillation to integrate capabilities from domain-specialized teachers. To better meet the low-latency requirements of on-device scenarios, IronLLM-0.6B adopts an Instruct-Only design. Evaluations show that IronLLM-0.6B achieves competitive performance relative to larger models such as Qwen3.5-0.8B and MiniCPM5-1B, while producing more concise responses on many tasks. We further present IronLLM-0.6B-Light, which replaces RMSNorm with Dynamic Tanh and simplifies several computationally expensive components to improve inference and quantization efficiency. Together, the IronLLM models provide an effective performance-efficiency trade-off for resource-constrained deployment.

CommentsTechnical report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑