arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28554cs.AIcs.CV

Pistis 技术报告

Pistis Technical Report

发表机构字节跳动
查看机构详情
  • ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

Heyun Chen, Xiaohan Lan, Jiaxi Li, Zhilin Lu, Qi She, Weiwen Xu, Fei Yu, Yujie Zhong, Jinghuan Chen, Zijian Feng, Siyu Jiao, Yiheng Lin, Xinhao Wang, Sihan Yang… 展开作者

Heyun Chen, Xiaohan Lan, Jiaxi Li, Zhilin Lu, Qi She, Weiwen Xu, Fei Yu, Yujie Zhong, Jinghuan Chen, Zijian Feng, Siyu Jiao, Yiheng Lin, Xinhao Wang, Sihan Yang, Jieyu You, Changbin Zhang, Hengyu Zhang, Xudong Zhang, Yunqing Zhao, Shuai Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Pistis模型系列及交错蒸馏与强化学习(IDRL)后训练框架,并引入系统级方法PAH,在不更新参数和增加交互预算下提升智能体性能。

中文摘要 AI 辅助

我们推出了 Pistis 模型系列,包含 27B 和 9B 参数的多模态大语言模型,分别基于 Qwen3.6 和 Qwen3.5 构建,并通过一个通用且可扩展的后训练框架开发。该框架首先通过大规模多模态监督微调(SFT)建立坚实基础。在此 SFT 基础之上,我们提出了交错蒸馏与强化学习(IDRL),这是一种新颖的后训练范式,在单个训练循环内紧密整合了在线策略蒸馏和强化学习。通过在两个目标之间交替进行,而不是单独优化任一目标或将其组合成静态联合损失,IDRL 实现了更有效的知识迁移、更高的优化稳定性,以及对于长程智能体轨迹更精确的信用分配,从而在缓解常见能力权衡的同时带来更强的性能。在两种模型规模下,该框架均产生两个专业变体:Pistis-Thinking,旨在增强深度多模态推理;以及 Pistis-Agentic,额外整合了智能体轨迹数据以支持长程规划、迭代推理和工具使用。Pistis-Agentic 在多模态搜索方面尤为强大。两种规模均优于其对应的基础模型。除模型参数优化外,我们进一步引入了 Pistis-Auto-Harnessing(PAH),一种系统级方法,通过迭代优化自动改进智能体的推理框架。实验表明,PAH 在不更新模型参数或增加交互预算的情况下提升了模型性能。

英文摘要

We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long-horizon agentic trajectories, leading to stronger performance while mitigating common capability trade-offs. At both model scales, the framework produces two specialized variants: Pistis-Thinking, designed to strengthen deep multimodal reasoning, and Pistis-Agentic, which additionally incorporates agentic trajectory data to support long-horizon planning, iterative reasoning, and tool use. Pistis-Agentic is particularly strong in multimodal search. Both scales outperform their corresponding base models. Beyond model-parameter optimization, we further introduce Pistis-Auto-Harnessing (PAH), a system-level method that automatically improves the agent's inference harness through iterative optimization. Experiments demonstrate that PAH enhances the model performance without updating the model parameters or increasing the interaction budget.

↑