AI 中文总结
研究针对具身智能体对设备端紧凑模型需求,提出雅典娜大脑8B模型,经多阶段训练流程,使其兼具通用与具身交互能力,实验证明该模型在通用和具身评估中有效,能在保持性能同时生成更简洁回复。
AI 中文摘要
大语言模型在语言理解、推理和世界知识方面展现出卓越能力。随着具身智能体能力增强,对能作为设备端大脑的紧凑模型需求渐长。现有方法常顾此失彼。本文提出雅典娜大脑8B模型,经多阶段训练流程,它兼具强大通用能力与具身交互能力,能生成简洁回复。实验表明其在通用和具身评估中均有效,相比Qwen3 - 8B思维模型,在通用语言和推理基准上性能相当但回复更短,在具身基准中表现出色。
英文摘要
Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, there is a growing demand for compact models that can serve as an on-device brain, preserving the broad general intelligence of LLMs while enabling effective high-level interaction with embodied environments. Existing approaches, however, often prioritize either general-purpose intelligence or specialized embodied capabilities, making it challenging to satisfy both requirements within a single model. We present \textbf{Athena-Brain-8B}, an 8B LLM designed to serve as an on-device brain for embodied intelligence for embodied intelligence. Through a multi-stage post-training pipeline consisting of General Supervised Fine-Tuning, General Reinforcement Learning, Embodied Expert training, and Model Merge, Athena-Brain-8B maintains strong general capabilities while acquiring strong high-level embodied interaction capabilities and generating concise responses for efficient embodied interaction. Experimental results demonstrate the effectiveness of Athena across both general and embodied evaluations. Compared with the corresponding Qwen3-8B thinking model, Athena-Brain-8B achieves comparable performance on general language and reasoning benchmarks while generating substantially shorter responses. On in-domain embodied benchmarks, Athena-Brain-8B consistently outperforms models of similar scale and surpasses several substantially larger frontier models evaluated zero-shot, demonstrating that compact language models can effectively integrate strong general intelligence with embodied capabilities.