Pistis 技术报告
Pistis Technical Report
查看机构详情
- ByteDance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出Pistis模型系列及交错蒸馏与强化学习(IDRL)后训练框架,并引入系统级方法PAH,在不更新参数和增加交互预算下提升智能体性能。
中文摘要 AI 辅助
我们推出了 Pistis 模型系列,包含 27B 和 9B 参数的多模态大语言模型,分别基于 Qwen3.6 和 Qwen3.5 构建,并通过一个通用且可扩展的后训练框架开发。该框架首先通过大规模多模态监督微调(SFT)建立坚实基础。在此 SFT 基础之上,我们提出了交错蒸馏与强化学习(IDRL),这是一种新颖的后训练范式,在单个训练循环内紧密整合了在线策略蒸馏和强化学习。通过在两个目标之间交替进行,而不是单独优化任一目标或将其组合成静态联合损失,IDRL 实现了更有效的知识迁移、更高的优化稳定性,以及对于长程智能体轨迹更精确的信用分配,从而在缓解常见能力权衡的同时带来更强的性能。在两种模型规模下,该框架均产生两个专业变体:Pistis-Thinking,旨在增强深度多模态推理;以及 Pistis-Agentic,额外整合了智能体轨迹数据以支持长程规划、迭代推理和工具使用。Pistis-Agentic 在多模态搜索方面尤为强大。两种规模均优于其对应的基础模型。除模型参数优化外,我们进一步引入了 Pistis-Auto-Harnessing(PAH),一种系统级方法,通过迭代优化自动改进智能体的推理框架。实验表明,PAH 在不更新模型参数或增加交互预算的情况下提升了模型性能。
英文摘要
We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long-horizon agentic trajectories, leading to stronger performance while mitigating common capability trade-offs. At both model scales, the framework produces two specialized variants: Pistis-Thinking, designed to strengthen deep multimodal reasoning, and Pistis-Agentic, which additionally incorporates agentic trajectory data to support long-horizon planning, iterative reasoning, and tool use. Pistis-Agentic is particularly strong in multimodal search. Both scales outperform their corresponding base models. Beyond model-parameter optimization, we further introduce Pistis-Auto-Harnessing (PAH), a system-level method that automatically improves the agent's inference harness through iterative optimization. Experiments demonstrate that PAH enhances the model performance without updating the model parameters or increasing the interaction budget.