Aero Realtime:用于低延迟流式多模态生成的完全对齐输入输出流
Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation
浏览论文内容
中文总结 AI 辅助
本文提出4B参数流式多模态模型Aero Realtime,采用双工架构实现输入输出流完全对齐,通过时隙对齐等技术降低延迟,在4块NVIDIA A6000 GPU上验证了其低延迟多模态交互的可行性。
中文摘要 AI 辅助
现有流式多模态模型虽增量式处理观测结果,但仍遵循“预填充后解码”的回合制模式,属于非双工架构:新观测无法自然接入活跃生成流。主动式替代方案采用微回合轮询或外部响应门,会割裂连续交互、使响应时机与语言生成解耦,还会增加KV缓存友好型服务的复杂度。本文提出Aero Realtime,这是一款具备双工架构的4B参数流式多模态模型,用于实时生成。Aero Realtime将视频、音频和文本输出对齐至共享时间网格,每个约80ms的音频时隙可预测词汇 token 或静音 token,这使输入与输出能同步推进,让单一自回归目标同时学习响应时机与生成内容。推理阶段,Aero Realtime仅追加最新多模态时隙、继承前序输出状态并复用KV缓存以实现高效增量执行。本文还提供完整的训练与服务方案,包括实时问答构建、时隙对齐监督、硬件感知分布式训练及可恢复推理。在4块NVIDIA A6000工作站GPU上,Aero Realtime在20分钟连续流视频处理中,保持84ms中位数和173ms P95处理延迟,始终处于源时间线的200ms范围内。这些结果证明,完全对齐输入输出建模可实现双工、主动且硬件适配的多模态交互。
英文摘要
Existing streaming multimodal models process observations incrementally but still follow a turn-based prefill-then-decode pattern, making them non-duplex: new observations cannot naturally enter an active generation stream. Proactive alternatives use micro-turn polling or external response gates, which fragment continuous interaction, decouple response timing from language generation, and complicate KV-cache-friendly serving. We introduce Aero Realtime, a 4B streaming multimodal model with a duplex architecture for realtime generation. Aero Realtime aligns video, audio, and textual output on a shared temporal grid, where each approximately 80-ms audio slot predicts either a lexical token or a silence token. This allows input and output to advance together, enabling one autoregressive objective to learn both when to respond and what to generate. During inference, Aero Realtime appends only the newest multimodal slot, carries forward the previous output state, and reuses the KV cache for efficient incremental execution. We further provide a complete training and serving recipe, including realtime QA construction, slot-aligned supervision, hardware-aware distributed training, and resumable inference. On four NVIDIA A6000 workstation GPUs, Aero Realtime maintains 84-ms median and 173-ms P95 processing lag over 20 minutes of a continuously streamed video, remaining within 200~ms of the source timeline. These results demonstrate the feasibility of fully aligned input-output modeling for duplex, proactive, and hardware-aligned multimodal interaction.