发表机构
Taobao & Tmall Group of Alibaba(阿里巴巴淘宝天猫集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出专为电商直播设计的全模态理解模型TLive-Omni,通过引入Per-vGrid等技术,结合多阶段训练流程,在电商直播基准任务上表现强劲且泛化能力出色。
AI 中文摘要
电商直播需要对嘈杂、时间跨度长的流进行全模态理解,产品信息分布在语音、视频帧、产品图像、叠加文本和用户查询中。我们提出TLive-Omni,这是一种专为直播电商场景设计的全模态理解模型,它将图像、视频、音频和文本输入映射到统一表示空间。针对长视频直播分析,我们引入Per-vGrid,这是一种带时间戳的令牌组织方式,将每个视频网格与其时间对应的音频分组,并通过显式边界令牌促进时间对齐。我们设计了一个三阶段监督训练流程,逐步发展直播电商理解能力,从全模态感知到指令跟随响应。随后我们提出Faithful-RFT,这是一种强化微调阶段,可进一步提升答案的忠实度和表达质量,同时满足实时需求,该阶段通过任务可验证反馈直接评分最终响应,而非在推理过程中优化推理式探索。此外,TLive-Omni得到面向场景的原子能力分类法和紧凑型数据生成引擎的支持,该引擎将直播电商的音频、图像和视频流转换为语音识别、说话人分析、产品视觉 grounding、文本识别、时间 grounding、视频密集字幕和全模态问答等任务的训练信号。为实现可扩展训练,同步长度分组采样器减少填充,同时保持各工作节点的工作量相当,而轻量动态采样策略则以接近零的奖励方差重新生成推理组,为GRPO维持有意义的相对优势。在电商直播基准上的实验表明,该模型在直播电商领域任务中表现强劲,同时在通用基准上也具备出色的泛化能力。
英文摘要
E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We present TLive-Omni, an omni-modal understanding model tailored to live-commerce scenarios. It maps image, video, audio, and text inputs into a unified representation space. For long-form live streaming analysis, we introduce Per-vGrid, a timestamped token organization that groups each video grid with its temporally corresponding audio within explicit boundary tokens to facilitate temporal alignment. We design a three-stage supervised training recipe that progressively develops live-commerce understanding, from omni-modal perception to instruction-following responses. We then propose Faithful-RFT, a reinforcement fine-tuning stage that further improves answer faithfulness and expression quality while meeting real-time demands, scoring final responses directly with task-verifiable feedback rather than optimizing for reasoning-style exploration during rollout. Moreover, TLive-Omni is supported by a scenario-oriented atomic capability taxonomy and a compact data production engine that converts live-commerce audio, image, and video streams into training signals for speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc. For scalable training, a synchronized length-grouped sampler reduces padding while preserving comparable workloads across workers, while a lightweight dynamic sampling strategy regenerates rollout groups with near-zero reward variance to maintain meaningful relative advantages for GRPO. Experiments on e-commerce live streaming benchmarks demonstrate strong performance across live-commerce domain tasks, together with excellent generalization on general benchmarks.