arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

Alibaba(阿里巴巴)

至 收录 1482
2607.17977 2026-07-21 cs.RO 新提交

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

RynnBrain 1.1:迈向更具能力和通用性的具身基础模型

Kehan Li, Bohan Hou, Minghao Zhu, Tianyi Zhang, Zesen Cheng, Zhikai Wang, Sicong Leng, Xin Li, Xiao Lin, Biying Yao, Minghua Zeng, Jiangpin Liu, Ronghao Dang, Jiayan Guo, Siteng Huang, Haoyu Zhao, Heng Ping, Yaxi Zhao, Kexiang Wang, Tong Lu, Shengke Xue, Jiahao Tang, Yulei Wang, Zejing Wang, Jianwei Gao, Shijian Lu, Chengju Liu, Jianfei Yang, Mingxiu Chen, Deli Zhao

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(湖畔实验室)

AI总结 介绍RynnBrain 1.1具身基础模型家族,其经统一框架训练,相比1.0版本有改进。开发RynnBrain - VLA并部署。该模型在多方面成果显著,初始化策略表现优,联合训练提升了得分和成功率。

Comments KL,BH,MZ,TZ,ZC,ZW,SL,XL,XL,BY,MZ,JL,RD contribute equally. Project Lead: Kehan Li and Xin Li project: https://alibaba-damo-academy.github.io/RynnBrain github: https://github.com/alibaba-damo-academy/RynnBrain huggingface: https://huggingface.co/collections/Alibaba-DAMO-Academy/rynnbrain-11 modelscope: https://modelscope.cn/collections/DAMO_Academy/RynnBrain-11

详情
AI中文摘要

我们展示了RynnBrain 1.1,这是一个涵盖2B、9B和122B - A10B规模的具身基础模型家族。它通过统一的时空和物理基础框架进行训练,支持具身感知、空间推理、定位和规划。与RynnBrain 1.0相比,它在整个模型家族中进一步引入了接触点预测,并为2B和9B模型引入了原生3D基础。我们还开发了具有统一跨具身动作空间和特定具身掩码的RynnBrain - VLA,并将其部署在Unitree G1、Astribot - S1和Tianji - Wuji上。RynnBrain 1.1在具身认知、定位和3D基础方面取得了显著成果,122B - A10B模型在VSI - Bench、MMSI和RefSpatial - Bench上优于所有评估的专有和开源模型。真实机器人实验表明,以RynnBrain初始化的策略优于基于Qwen的和有代表性的通用VLA,而联合多任务和多具身训练比单任务训练提高了过程得分和成功率。

英文摘要

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.

URL PDF HTML 收藏
2607.17779 2026-07-21 cs.AI 新提交

Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

动态防御剖析实现文本到图像模型的认知越狱

Dongdong Yang, Deyue Zhang, Zhao Liu, Zonghao Ying, Wenzhuo Xu, Jiankai Jin, Xiangzheng Zhang, Quanchen Zou

机构 * Alibaba Tongyi Lab(阿里巴巴通义实验室)

AI总结 研究文本到图像模型易受对抗性滥用问题,提出MIND认知越狱框架,将对抗提示生成重构为信念状态推理,集成多模态判断器等组件,经实验验证其在多个防御设置下显著优于现有方法,有效实现越狱生成。

Comments 12 pages, 3 figures

详情
AI中文摘要

文本到图像(T2I)生成模型在合成高质量视觉内容方面取得了显著进展,但仍容易受到对抗性滥用,特别是在生成不适宜工作的(NSFW)图像方面。大多数现有越狱攻击主要依赖启发式提示工程或黑箱优化,将模型反馈视为二元信号(成功或失败)。这种粗粒度范式忽略了各种失败模式中嵌入的丰富信息,导致探索效率低下和严重的语义崩溃。本文提出了MIND,一个认知越狱框架,将对抗性提示生成重新构建为对潜在防御机制的信念状态推理问题。MIND通过将多模态反馈解释为高密度信号,积极地对目标系统的潜在防御机制进行建模。具体来说,该框架集成了三个核心组件:(1)用于细粒度反馈分解的多模态判断器,(2)用于迭代信念更新的防御剖析器,以及(3)用于检索历史有效攻击策略的元记忆模块。这些组件在一个推理驱动的进化优化过程中统一起来,实现自适应和语义一致的越狱生成。在I2P基准上的大量实验证明了MIND的有效性。在应用于Stable Diffusion v1.5模型的六种代表性预处理和后处理防御设置下,MIND实现了95.62%的攻击成功率(ASR),显著优于现有方法。此外,所提出框架的有效性在四个广泛使用的商业T2I系统中得到验证,在Wan-2.5上实现了91.58%的最高ASR。

英文摘要

Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulnerable to adversarial misuse, particularly in generating Not-Safe-For-Work (NSFW) images. Most existing jailbreak attacks primarily rely on heuristic prompt engineering or black-box optimization, treating model feedback as a binary signal (success or failure). This coarse-grained paradigm overlooks the rich information embedded in diverse failure modes, such as textual refusal, visual blocking, and semantic sanitization, resulting in inefficient exploration and severe semantic collapse. In this paper, we propose MIND, a cognitive jailbreak framework that reframes adversarial prompt generation as a belief-state inference problem over latent defense mechanisms. Instead of blindly searching for bypass prompts, MIND actively models the target system's latent defense mechanisms by interpreting multi-modal feedback as high-density signals. Specifically, the framework integrates three core components: (1) a Multi-modal Judge for fine-grained feedback decomposition, (2) a Defense Profiler for iterative belief updating, and (3) a Meta-Memory module for retrieving historically effective attack strategies. These components are unified within a reasoning-driven evolutionary optimization process, enabling adaptive and semantically consistent jailbreak generation. Extensive experiments on the I2P benchmark demonstrate the effectiveness of MIND. Under six representative pre-processing and post-processing defense settings applied to the Stable Diffusion v1.5 model, MIND achieves an Attack Success Rate (ASR) of 95.62%, significantly outperforming existing methods. Additionally, the effectiveness of the proposed framework is validated across four widely used commercial T2I systems, achieving the highest ASR of 91.58% on Wan-2.5.

URL PDF HTML 收藏
2607.17715 2026-07-21 cs.CL 新提交

C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference

C$^2$KV:用于高效大语言模型推理的压缩可组合键值缓存重用

Chuheng Du, Junyi Chen, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Chaoyue Niu, Shengzhong Liu, Guihai Chen, Fan Wu

机构 * Alibaba Group(阿里巴巴集团)

AI总结 研究针对长上下文LLM推理成本高问题,提出C$^2$KV框架,联合优化KV提取与推理拼接,学习可组合压缩的KV缓存流形,引入轻量级边车提取器及协同训练策略,显著降低缓存成本,实现推理加速并保持生成质量。

Comments 12 pages, 9 figures, accepted by ACM SIGKDD 2026

详情
AI中文摘要

长上下文推理是现代大语言模型(LLM)应用(如检索增强生成和多文档推理)的核心。为减轻不断增长的推理成本,近期工作探索了键值(KV)缓存重用以减少冗余预填充计算。但现有方法主要关注计算节省,忽视了长上下文LLM服务中的关键瓶颈:存储和访问大型KV缓存的成本。虽然KV压缩似乎是自然补充,但将其与非前缀KV重用简单结合常导致严重的精度下降。在这项工作中,我们提出C$^2$KV,一个用于非前缀KV重用的统一框架,联合优化KV提取和推理时的拼接。C$^2$KV学习一个可组合且压缩的KV缓存流形,明确设计为与位置无关。我们的方法引入了一个轻量级的带有可学习压缩令牌和结构化注意力流的边车提取器,实现模块化的KV表示,可灵活重用和拼接而无需修改冻结的基础模型。我们还采用了压缩 - 拼接协同训练策略,使提取时的表示与其下游重用行为对齐。跨多个长上下文基准和模型系列的广泛实验表明,C$^2$KV显著降低了KV缓存存储和传输成本,在长上下文下实现高达17倍的推理加速,同时保持生成质量。

英文摘要

Long-context inference is central to modern large language model (LLM) applications such as retrieval-augmented generation and multi-document reasoning. To mitigate the growing inference cost, recent work has explored key-value (KV) cache reuse to reduce redundant prefill computation. However, existing reuse methods primarily focus on computation savings and overlook a critical bottleneck in long-context LLM serving: the cost of storing and accessing large KV caches. While KV compression appears to be a natural complement, naively combining compression with non-prefix KV reuse often leads to severe accuracy degradation. In this work, we propose C$^2$KV, a unified framework for non-prefix KV reuse that jointly optimizes KV extraction and inference-time concatenation. C$^2$KV learns a composable and compressed KV cache manifold that is explicitly designed to be position-agnostic. Our approach introduces a lightweight sidecar Extractor with learnable compression tokens and a structured attention flow, enabling modular KV representations that can be flexibly reused and concatenated without modifying the frozen base model. We further employ a compression-concatenation co-training strategy to align extraction-time representations with their downstream reuse behavior. Extensive experiments across multiple long-context benchmarks and model families demonstrate that C$^2$KV significantly reduces KV cache storage and transfer costs, achieving up to 17$\times$ inference speedup under long contexts, while preserving generation quality.

URL PDF HTML 收藏
2607.17499 2026-07-21 cs.AI 新提交

Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation

派利淘-多模态搜索:构建原生电子商务多模态搜索基础

Xiaohan Ye, Xu Chen, Zihan Gong, Jian Ding, Lianyu Du, Baicheng Chen, Yunmeng Shu, Jingqian Zhao, Zhixiang Zhao, Shuaiqi Jia, Chong Ma, Shuwen Xiao, Xiangheng Kong, Yuan Gao, Jun Song, Jinsong Lan, Xiaoyong Zhu, Bo Zheng

机构 * Taobao & Tmall Group of Alibaba(阿里巴巴淘宝及天猫集团)

AI总结 针对电子商务多模态搜索问题,提出派利淘-多模态搜索基础模型,通过混合语义ID、两阶段持续预训练策略和混合推理后训练管道,在淘宝派利淘平台实现显著改进,提升了商品交易总额和交易量。

Comments Technical Report: Pailitao-MMSearch

详情
AI中文摘要

电子商务的发展使产品搜索从简单文本关键词查询转变为复杂多模态交互。现有方法面临困境:单模态专家模型孤立运行无法处理跨模态查询,通用视觉语言模型缺乏领域特定知识。本文提出派利淘-多模态搜索基础模型,引入三项关键创新:混合语义ID、两阶段持续预训练策略和混合推理后训练管道。基于文生模型构建并部署在淘宝派利淘平台,在线A/B测试取得显著改进,证明了该模型的有效性。

英文摘要

The evolution of e-commerce has fundamentally transformed how users search for products, shifting from simple text-based keyword queries to complex multimodal interactions that seamlessly combine product images, natural language descriptions, and mixed-intent instructions. However, existing approaches face a critical dilemma: single-modal specialist models, deployed independently for text retrieval, visual search, and voice recognition, operate in isolation and cannot handle cross-modal queries, while general-purpose vision-language models lack the domain-specific knowledge necessary for fine-grained product understanding, user behavior modeling, and commercial intent reasoning. In this work, we present Pailitao-MMSearch, one native e-commerce multimodal search foundation model designed to bridge this gap. Our approach introduces three key innovations: (1)HybSID (Hybrid Semantic ID);(2)a two-stage continual pre-training strategy; and (3)a hybrid reasoning post-training pipeline. Built upon Qwen and deployed on Taobao's Pailitao multimodal search platform, Pailitao-MMSearch achieves substantial improvements in online A/B testing, including up to +13.61\% in Gross Merchandise Volume (GMV) and +8.21\% in transaction volume compared to traditional multi-modal search pipeline, demonstrating the effectiveness of our native e-commerce multimodal search large language models.

URL PDF HTML 收藏
2607.17281 2026-07-21 cs.LG cs.AI 新提交

AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization

AIGB-R1:通过分层规划器-执行器优化实现自我进化的生成式自动出价

Yuejia Dou, Hesong Wang, Xinyu Zhang, Tianyu Wang, Zhilin Zhang, Chuan Yu, Jian Xu, Bo Zheng, Qi Qi

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院) Alibaba Group(阿里巴巴集团)

AI总结 研究针对AIGB范式在自动出价中存在的问题,提出AIGB-R1框架,利用大语言模型推理能力,通过分层规划器与执行器模块、经验驱动循环、两阶段训练及新优化方法,经实验验证该框架在自动出价任务中的有效性。

详情
AI中文摘要

自动出价在在线广告中起着至关重要的作用,能自动调整出价以优化广告商的商业目标。新兴的人工智能生成出价(AIGB)范式广泛采用生成模型来优化出价策略,但存在离线数据集模式覆盖有限和任务状态理解不足的问题,阻碍了对最优策略的有效探索。大语言模型(LLMs)具有先验世界知识和推理能力,为克服这些限制提供了有前景的方法。然而,直接将LLMs应用于自动出价任务面临数值精度有限、幻觉和推理延迟等固有挑战。为解决这些限制,我们提出了AIGB-R1,这是一个分层的自我进化自动出价框架,旨在通过LLMs的推理能力增强人工智能生成出价,包括用于宏观策略规划的高级规划器模块和用于细粒度决策的低级执行器模块。在此基础上,我们设计了一个经验驱动的自我进化循环,从积累的经验中实现自主策略探索和优化。我们采用离线预训练和训练后对齐的两阶段管道,并构建了一个用于策略展开的交互式出价模拟环境。此外,我们提出了解耦组相对策略优化(D-GRPO),通过优势解耦实现端到端优化。在大规模公共数据集上的实验结果证明了AIGB-R1的有效性。

英文摘要

Auto-bidding plays an essential role in online advertising, automatically adjusting bids for advertisers to optimize their commercial goals. The emerging AI-Generated Bidding (AIGB) paradigm widely adopts generative modeling to optimize bidding strategies, yet suffers from the limited mode coverage of offline datasets and inadequate task-state understanding, hindering effective exploration of optimal strategies. Large Language Models (LLMs), with prior world knowledge and reasoning capabilities, offer a promising approach to overcome these limitations. However, directly applying LLMs to auto-bidding tasks faces inherent challenges in limited numerical precision, hallucinations, and inference latency. To address these limitations, we propose AIGB-R1, a hierarchical self-evolving auto-bidding framework aiming to enhance AI-Generated Bidding via LLMs' Reasoning capabilities, comprising a high-level Planner module for macro-level strategy planning and a low-level Executor module for fine-grained decision-making. Building upon this, we design an experience-driven self-evolving loop, enabling autonomous strategy exploration and optimization from accumulated experience. We adopt a two-stage pipeline of offline pre-training and post-training alignment, and build an interactive bidding simulation environment for strategy rollout. Furthermore, we propose Decoupled Group Relative Policy Optimization (D-GRPO) to achieve end-to-end optimization via advantage decoupling. Experimental results on a large-scale public dataset demonstrate the effectiveness of AIGB-R1.

URL PDF HTML 收藏
2607.16828 2026-07-21 cs.CV 新提交

UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation

UniNDM:文本到图像生成中针对性内容的统一噪声驱动检测与缓解框架

Yao Huang, Yitong Sun, Huanran Chen, Ruochen Zhang, Shouwei Ruan, Ranjie Duan, Maoxun Yuan, Yinpeng Dong, Hui Xue, Xiaochun Cao, Xingxing Wei

机构 * Institute of Artificial Intelligence, State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室人工智能研究院) College of Artificial Intelligence, Tsinghua University(清华大学人工智能学院) Security Department, Alibaba Group(阿里巴巴集团安全部) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-Sen University(中山大学深圳校区网络科学与技术学院)

AI总结 针对文本到图像生成易受隐式性提示影响的问题,提出UniNDM统一噪声驱动框架。利用早期预测噪声的可分离性开发轻量级检测器,引入噪声增强自适应负引导缓解问题,扩展到扩散变压器架构,实验显示比现有方法有显著改进。

Comments 18 pages, 10 figures, accepted by TPAMI

详情
AI中文摘要

尽管文本到图像扩散模型具有强大的生成能力,但它们容易受到隐式性提示的影响,由于模型偏差或训练数据中的潜在相关性,微妙线索会意外生成不当内容。现有安全机制存在根本局限性。为此,我们提出UniNDM,一个统一的噪声驱动框架,通过扩散过程中的噪声动态来重新思考安全机制。我们发现早期预测噪声在正常和性明确内容之间具有内在可分离性,并理论证明其语义浓度随时间步长二次增加。利用此特性,我们开发了轻量级基于噪声的检测器,准确率高且几乎无计算开销。对于缓解,我们引入噪声增强自适应负引导,通过大语言模型动态生成特定上下文负提示,同时通过抑制对明确令牌的注意力集中来优化初始噪声。我们还将框架扩展到新兴的扩散变压器架构。综合实验表明,我们的方法比现有方法有显著改进。

英文摘要

Despite the impressive generative capabilities of text-to-image diffusion models, they remain vulnerable to implicit sexual prompts, where subtle cues disguised as benign terms or adversarial tokens unexpectedly generate the inappropriate content due to model biases or latent correlations in training data. Existing safety mechanisms face fundamental limitations: detection methods primarily identify explicit content and fail to capture implicit malicious intent, while mitigation approaches rely on static negative prompts inadequate for diverse implicit scenarios. To address these challenges, we propose UniNDM, a unified noise-driven framework that rethinks safety mechanisms through the lens of noise dynamics in diffusion processes. Our key insight is that early-stage predicted noise exhibits inherent separability between normal and sexually explicit content, which we theoretically demonstrates quadratically increasing semantic concentration with timestep. Leveraging this property, we develop a lightweight noise-based detector achieving superior accuracy with virtually no computational overhead. For mitigation, we introduce noise-enhanced adaptive negative guidance: dynamically generating context-specific negative prompts via large language models to handle diverse implicit content, while optimizing initial noise by suppressing attention concentration on explicit tokens to provide comprehensive protection. Besides the U-Net-based diffusion models, we further extend our framework to emerging Diffusion Transformer architectures through region-constrained semantic guidance tailored for their unified multimodal attention. Comprehensive experiments across U-Net models and DiT models on both natural and adversarial datasets demonstrate substantial improvements over state-of-the-art methods, including SLD, UCE, Safree, etc. Our code is publicly available at https://github.com/Aries-iai/UniNDM.

URL PDF HTML 收藏
2607.16692 2026-07-21 cs.SE cs.CL 新提交

Dependency-Guided Code Generation: Structured Matrix Decomposition and Consistency-Guided Refinement

依赖引导的代码生成:结构化矩阵分解与一致性引导的细化

Mingqiao Mo, Yangchen Zeng, Zikai Xiao, Xin Xiao, Wenhua Nie, Zhaolu Kang, Guangyuan Dong, Kai Shu, Hao Zhang, Xiaodong Fan

机构 * University of the Chinese Academy of Sciences(中国科学院大学) ByteDance Inc.(字节跳动公司) Zhejiang University(浙江大学) National Taiwan University(台湾国立大学) Peking University(北京大学) Alibaba Group(阿里巴巴集团) Tsinghua University(清华大学) Liaoning Technical University(辽宁技术大学)

AI总结 针对现有代码生成方法无法充分捕捉代码实体依赖关系的问题,提出依赖感知代码生成框架,通过结构化矩阵分解和一致性引导细化生成代码,经实验验证该方法能生成语义对齐和结构保真度更高的代码。

Comments 12 pages

详情
AI中文摘要

现代软件系统日益复杂,使自动代码生成成为软件工程中的一项基本任务。然而,现有方法往往无法充分捕捉代码实体间复杂的多层次依赖关系,导致生成的代码逻辑不完整或难以集成到实际系统中。为解决此局限,我们提出一个依赖感知代码生成框架,通过基于图的表示明确建模代码实体间的交互。我们将依赖分解为两个互补组件:一个捕获强显式关系的量化矩阵和一个对弱隐式交互建模的稀疏低秩分解。通过交替优化过程有效学习分解。在代码生成期间,将学习到的依赖结构作为约束纳入,确保生成代码的语义连贯和结构一致。此外,我们为强依赖引入稀疏三元组表示,显著提高存储效率和计算可扩展性。大量实验表明,与现有方法相比,我们的方法始终能生成具有更高语义对齐和结构保真度的代码。

英文摘要

The increasing complexity of modern software systems has made automated code generation a fundamental task in software engineering. However, existing approaches often fail to adequately capture the intricate, multi-level dependencies among code entities, leading to generated code that is logically incomplete or difficult to integrate into real-world systems. To address this limitation, we propose a dependency-aware code generation framework that explicitly models interactions among code entities through a graph-based representation. We decompose dependencies into two complementary components: a quantized matrix that captures strong, explicit relations, and a sparse low-rank factorization that models weaker, implicit interactions. The decomposition is efficiently learned via an alternating optimization procedure. During code generation, the learned dependency structure is incorporated as a constraint, ensuring both semantic coherence and structural consistency of the generated code. Furthermore, we introduce a sparse triplet representation for strong dependencies, significantly improving storage efficiency and computational scalability. Extensive experiments demonstrate that our approach consistently produces code with superior semantic alignment and structural fidelity compared to existing methods.

URL PDF HTML 收藏
2607.16673 2026-07-21 cs.CL 新提交

SpecLA: Efficient Speculative Decoding for Linear-Attention Models

SpecLA:线性注意力模型的高效推测解码

Zhibin Wang, Xuying Han, Zhaohua Yang, Fuliang Liu, Xue Li, Rong Gu, Sheng Zhong, Chen Tian

机构 * Alibaba Group(阿里巴巴集团)

AI总结 研究针对线性注意力模型自回归解码效率低的问题,提出SpecLA推测解码运行时,通过拓扑感知内核验证、存储紧凑因子及置信度修剪等方法,在NVIDIA H100上实现了比自回归解码高达1.70倍的端到端加速。

详情
AI中文摘要

线性注意力模型用循环状态取代不断增长的KV缓存,但自回归解码仍一次一个令牌地读取、更新和写入这些状态。推测解码可通过在一次目标传递中验证多个草稿令牌来降低成本,但现有推测系统是为Transformer KV缓存设计的。对于有状态线性注意力目标,验证必须遵循跨链和分支的循环依赖,接受必须仅更新接受的状态轨迹,起草者必须避免提交浪费有状态验证工作的候选。本文提出SpecLA,一种有状态线性注意力模型的推测解码运行时。SpecLA用拓扑感知内核验证链和树,存储验证期间产生的紧凑因子以恢复接受状态,并使用置信度修剪和目标对齐的EAGLE风格起草者向验证器提供有用候选。在具有公共GDN-1.3B目标的NVIDIA H100上,SpecLA比自回归解码实现高达1.70倍的端到端加速。

英文摘要

Linear-attention models replace the growing KV cache with recurrent states, but autoregressive decoding still reads, updates, and writes these states one token at a time. Speculative decoding can reduce this cost by verifying several draft tokens in one target pass, yet existing speculative systems are designed for Transformer KV caches. For stateful linear-attention targets, verification must follow recurrent dependencies across chains and branches, acceptance must update only the accepted state trajectory, and the drafter must avoid submitting candidates that waste stateful verification work. This paper presents SpecLA, a speculative decoding runtime for stateful linear-attention models. SpecLA verifies chains and trees with topology-aware kernels, stores compact factors produced during verification to recover accepted states, and uses confidence pruning plus a target-aligned EAGLE-style drafter to feed useful candidates to the verifier. On an NVIDIA H100 with a public GDN-1.3B target, SpecLA achieves up to 1.70x end-to-end speedup over autoregressive decoding.

URL PDF HTML 收藏
2606.19341 2026-07-21 cs.CV cs.CL cs.SD 版本更新

Native Active Perception as Reasoning for Omni-Modal Understanding

原生主动感知作为全模态理解的推理

Zhenghao Xing, Ruiyang Xu, Yuxuan Wang, Jinzheng He, Ziyang Ma, Qize Yang, Yunfei Chu, Jin Xu, Junyang Lin, Chi-Wing Fu, Pheng-Ann Heng

机构 * The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Qwen Team, Alibaba Group(阿里巴巴集团Qwen团队)

AI总结 提出OmniAgent,一种基于POMDP迭代观察-思考-行动循环的原生全模态智能体,通过主动感知将推理复杂度与视频时长解耦,在多个基准上达到开源模型最优性能。

Comments Accepted at ICML 2026. Code and models: https://github.com/harryhsing/omniagent

详情
AI中文摘要

用于长视频理解的被动模型通常依赖于“全看一遍”范式,无论查询难度如何都统一处理帧,导致计算成本随视频时长增长。尽管出现了交互式框架,但它们通常依赖于全局预扫描,其上下文成本仍随视频长度扩展。我们提出OmniAgent,第一个原生全模态智能体,将视频理解建模为基于POMDP的迭代观察-思考-行动循环。OmniAgent执行按需动作,选择性地将视听线索提炼到持久文本记忆中,有效将推理复杂度与原始视频时长解耦。为实现这一点,我们引入了(1)智能体监督微调,通过最佳N轨迹合成和双阶段质量控制在启动原生主动感知;(2)带TAURA(轮次感知自适应不确定性重缩放优势)的智能体强化学习,利用轮次级熵将信用分配引导至关键发现轮次。关键的是,OmniAgent表现出正向测试时缩放,性能随推理轮次增加而提升,验证了主动感知的有效性。在十个基准(如VideoMME、LVBench)上的实验结果表明,OmniAgent在开源模型中达到了最先进性能。值得注意的是,在LVBench上,我们的7B智能体优于10倍大的Qwen2.5-VL-72B(50.5% vs. 47.3%)。

英文摘要

Passive models for long video understanding typically rely on a "watch-it-all" paradigm, processing frames uniformly regardless of query difficulty, causing computational cost to grow with video duration. Although interactive frameworks have emerged, they often rely on global pre-scanning, and their context cost still scales with video length. We propose OmniAgent, the first native omni-modal agent that formulates video understanding as a POMDP-based iterative Observation-Thought-Action cycle. OmniAgent executes on-demand actions to selectively distill audio-visual cues into a persistent textual memory, effectively decoupling reasoning complexity from raw video duration. To operationalize this, we introduce (1) Agentic Supervised Fine-Tuning to bootstrap native active perception via best-of-N trajectory synthesis with dual-stage quality control, and (2) Agentic Reinforcement Learning with TAURA (Turn-aware Adaptive Uncertainty Rescaled Advantage), which leverages turn-level entropy to steer credit assignment toward pivotal discovery turns. Crucially, OmniAgent exhibits positive test-time scaling, where performance improves as the number of reasoning turns increases, validating the efficacy of active perception. Empirical results across ten benchmarks (e.g., VideoMME, LVBench) demonstrate that OmniAgent achieves state-of-the-art performance among open-source models. Notably, on LVBench, our 7B agent outperforms the 10$\times$ larger Qwen2.5-VL-72B (50.5% vs. 47.3%).

URL PDF HTML 收藏
2603.22455 2026-07-21 cs.LG 版本更新

SkillRouter: Skill Routing for LLM Agents at Scale

SkillRouter:大规模LLM代理中的技能路由

YanZhao Zheng, ZhenTao Zhang, Chao Ma, YuanQiang Yu, JiHuai Zhu, Yong Wu, Tianze Xu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu

机构 * Alibaba Group(阿里巴巴集团)

AI总结 本文提出SkillRouter,通过全文本检索和重排序提升大规模LLM代理的技能路由性能,实验显示隐藏技能体导致路由准确率下降31-44个百分点,SkillRouter在基准测试中达到74.0%的Hit@1性能。

详情
AI中文摘要

可重用的技能使LLM代理能够将任务特定的程序、工具特性及执行指导封装成模块化组件。随着技能生态系统扩展到数万条目,推理时暴露每个技能变得不可行,从而产生技能路由问题:给定用户任务,系统必须在下游规划或执行前识别相关技能。现有代理堆栈通常依赖逐步披露,仅暴露技能名称和描述而隐藏完整实现体。我们在此基础上评估了这一设计选择,在一个基于SkillsBench的基准测试中,约有8万候选技能,针对实际重要的大规模技能注册表设置。在稀疏、密集和重排序基线中,隐藏技能体导致路由准确率下降31-44个百分点,显示完整技能文本是该设置中的关键路由信号而非次要元数据优化。受此发现启发,我们提出了SkillRouter,一个紧凑的12亿参数全文本检索和重排序流水线。SkillRouter在我们的基准测试中达到74.0%的Hit@1性能——在我们评估的基线中最强的平均Top-1路由性能——同时使用比最强基线流水线少13倍的参数,并运行速度快5.8倍。排名收益进一步泛化到一个独立构建的补充基准。在四个编码代理的互补端到端研究中,路由收益转移到任务成功度的提升,对于更强大的代理,收益更大。

英文摘要

Reusable skills let LLM agents package task-specific procedures, tool affordances, and execution guidance into modular building blocks. As skill ecosystems grow to tens of thousands of entries, exposing every skill at inference time becomes infeasible. This creates a skill-routing problem: given a user task, the system must identify relevant skills before downstream planning or execution. Existing agent stacks often rely on progressive disclosure, exposing only skill names and descriptions while hiding the full implementation body. We examine this design choice on a SkillsBench-derived benchmark with approximately 80K candidate skills, targeting the practically important setting of large skill registries with heavy overlap. Across representative dense and reranking baselines on this setting, hiding the skill body causes a 37-44 percentage point drop in routing accuracy. Stronger controls show that the missing signal is body-resident rather than a simple length artifact: body-distilled descriptions recover part of the gap, but remain 7-21 points below direct all-field routing, while a metadata-only encoder trained with the same data remains 14.0 points below its all-field counterpart. Motivated by this finding, we present Skillrouter, a compact 1.2B body-aware retrieve-and-rerank pipeline. Skillrouter achieves 74.0% Hit@1 on our benchmark -- the strongest average top-1 routing performance among the baselines we evaluate -- while using 13$\times$ fewer parameters and running 5.8$\times$ faster than the strongest base pipeline. The ranking gains further generalize to a supplementary benchmark independently constructed from three skill sources. In a complementary end-to-end study across four coding agents, routing gains transfer to improved task success, with larger gains for more capable agents.

URL PDF HTML 收藏
2510.27497 2026-07-21 cs.LG cs.AI 版本更新

InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames

InertialAR:基于惯性框架的自回归3D分子生成

Haorui Li, Weitao Du, Yuqiang Li, Hongyu Guo, Shengchao Liu

机构 * The Chinese University of Hong Kong(香港中文大学) Alibaba DAMO Academy(阿里巴巴达摩院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Ottawa(渥太华大学)

AI总结 研究探索基于Transformer的自回归模型在3D分子生成上的应用。提出InertialAR,通过规范标记化、几何位置编码及分层自回归范式应对挑战,在无条件和可控生成任务中表现出色,多个指标达领先水平。

Comments Accepted at ICML 2026

详情
AI中文摘要

基于Transformer的自回归模型已成为跨文本和图像等模态的统一范式,但其在3D分子生成方面的扩展仍未得到充分探索。这一差距源于两个基本挑战:一是如何将分子标记化为对SE(3)变换和原子索引排列均不变的规范1D标记序列;二是如何设计一种能够对将离散原子类型与连续3D坐标相结合的基于原子的混合标记进行建模的架构。为应对这些挑战,我们引入了InertialAR。它首先通过将每个分子与规范惯性框架对齐并重新排列原子来执行面向生成的规范标记化,将任意3D结构转换为用于自回归生成的唯一、SE(3)和排列不变的标记序列。在此规范标记化的基础上,我们提出了几何位置编码(GeoPE),赋予Transformer注意力3D几何感知能力。最后,InertialAR利用分层自回归范式解码下一个原子,通过扩散损失连续预测原子类型和3D坐标。实验表明,InertialAR在QM9、GEOM-Drugs和B3LYP的无条件生成的10个评估指标中的8个上取得了领先性能。此外,在可控生成以实现目标化学功能方面,它显著优于基线,在所有5个指标上均达到领先结果。代码可在指定网址获取。

英文摘要

Transformer-based autoregressive models have emerged as a unifying paradigm across modalities such as text and images, but their extension to 3D molecule generation remains underexplored. The gap stems from two fundamental challenges: (1) how to tokenize molecules into a canonical 1D sequence of tokens that is invariant to both SE(3) transformations and atom index permutations, and (2) how to design an architecture capable of modeling hybrid atom-based tokens that couple discrete atom types with continuous 3D coordinates. To address these challenges, we introduce InertialAR. It first performs generation-oriented canonical tokenization by aligning each molecule to a canonical inertial frame and reordering atoms, thereby converting arbitrary 3D structures into a unique, SE(3)- and permutation-invariant sequence of tokens for autoregressive generation. Built upon this canonical tokenization, we propose geometric positional encoding (GeoPE), which endows Transformer attention with 3D geometric awareness. Finally, InertialAR utilizes a hierarchical autoregressive paradigm to decode the next atom, consecutively predicting the atom type and 3D coordinates via Diffusion Loss. Experimentally, InertialAR achieves state-of-the-art performance on 8 of the 10 evaluation metrics for unconditional generation across QM9, GEOM-Drugs, and B3LYP. Moreover, it significantly outperforms baselines in controllable generation for targeted chemical functionality, attaining state-of-the-art results across all 5 metrics. Code is available at github.com/HaoruiLi46/InertialAR.

URL PDF HTML 收藏
2607.15810 2026-07-20 cs.LG 新提交

QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides

QUADS:通过双边量化误差对齐稳定用于专家混合模型的NVFP4强化学习

Zhengyang Zhuge, Hao Yu, Xin Wang, Zheng Li, Yizhong Cao, Dayiheng Liu, Jianwei Zhang

机构 * Alibaba Inc(阿里巴巴公司)

AI总结 研究针对MoE大语言模型RL中展开生成瓶颈,提出QUADS方法。通过训练 - 推理误差分析确定激活误差是FP4 RL不稳定主因,在训练器和展开侧分别采取措施,实现BF16精度,提升指标并提高展开吞吐量。

详情
AI中文摘要

在用于专家混合(MoE)大语言模型的强化学习(RL)中,展开生成是一个主要瓶颈,促使诸如FP8等低精度展开加速。作为一种新兴的低精度格式,NVFP4将用于精度保持的细粒度缩放与原生W4A4 FP4通用矩阵乘法相结合,以实现比FP8更高的吞吐量。然而,直接将NVFP4应用于MoE RL展开是不切实际的。NVFP4展开与BF16训练在大约150步后崩溃,伴随着展开 - 训练器对数概率差距的迅速扩大。通过训练 - 推理误差分析和控制消融,我们确定激活误差而非权重误差是FP4 RL不稳定的主要来源。为了稳定用于MoE的NVFP4 RL,我们提出了双边量化误差对齐(QUADS)。在训练器方面,我们引入非对称量化感知训练,对权重进行伪量化,同时保持激活不量化以实现更好的对齐。在展开方面,残差激活补偿在保留原生W4A4通用矩阵乘法的同时纠正高误差激活通道。在多个基准上的MoE RL实验中,QUADS实现了BF16级别的精度,比朴素的NVFP4 RL平均pass@值提高了21.49分,并且比FP8的展开吞吐量高约16%。

英文摘要

Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP8. As an emerging low-precision format, NVFP4 combines fine-grained scaling for accuracy preservation with native W4A4 FP4 GEMMs for higher throughput than FP8. However, we find that directly applying NVFP4 to MoE RL rollout is impractical. NVFP4 rollout with BF16 training collapses after roughly 150 steps, accompanied by rapidly growing rollout-trainer log-probability gaps. Through training-inference error analysis and controlled ablations, we identify activation error, rather than weight error, as the dominant source of FP4 RL instability: weights can be synchronized and aligned by a shared quantization-dequantization path, whereas activations are recomputed online and error is amplified by the coarse E2M1 grid. Therefore, to stabilize NVFP4 RL for MoE, we propose QUantization-error Alignment across Dual Sides (QUADS). On the trainer side, we introduce Asymmetric Quantization-Aware Training fake-quantizing weights while keeping activations unquantized for better alignment. On the rollout side, Residual Activation Compensation corrects high-error activation channels while preserving native W4A4 GEMMs. In our MoE RL experiments on several benchmarks, QUADS achieves BF16-level accuracy, improves average pass@1 by 21.49 points over naive NVFP4 RL, and delivers ~16% higher rollout throughput than FP8.

URL PDF HTML 收藏
2607.15766 2026-07-20 cs.CL 新提交

Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery

行动之前:在前瞻性假设发现任务上对大语言模型进行基准测试

Tianyun Zhong, Wangyi Jiang, Wei Wang, Xuanang Chen, Yaojie Lu, Shiwei Ye, Yuzhen Shi, Boyu Yang, Jinghang Wang, Han Li, Weiqi Zhai, Bing Zhao, Hu Wei, Haiyang Yu, Yongbin Li, Hongyu Lin, Le Sun, Xianpei Han

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) Alibaba Group(阿里巴巴集团)

AI总结 研究旨在评估大语言模型在开放式结论前阶段的发现能力,引入前瞻性假设发现任务及评估框架HypoArena,提出回顾性上下文回归构建基准数据,实验揭示模型能力分层及效应,结果支持将其作为评估大语言模型制定调查方向的独特目标。

详情
AI中文摘要

大语言模型在回答预先指定的问题方面表现出色,但其在开放式、结论前阶段的发现能力仍未得到充分衡量。我们引入了前瞻性假设发现(PHD),要求模型从不确定的证据(包括异常观测和碎片化记录)中自主构建有根据、有区分性且可测试的假设空间,以指导后续调查。为评估此能力,我们引入了HypoArena,包括HypoData(六个科学和分析领域的988个案例基准)和HypoEval(开放式假设集评估框架)。为大规模构建HypoData,我们提出回顾性上下文回归,这是一个通过去除明确结论、目标假设和回顾性因果归因同时保留事实基础,从完整专家文档重建结论前上下文的Forge - Audit管道。由于PHD允许多个有效输出,HypoEval结合双向成对判断与Bradley - Terry - Davidson聚合进行排名以及六维评分细则进行诊断。对15个前沿大语言模型的实验揭示了明显的能力分层和结构化分析技能的模型依赖效应,一些性能较低的模型在HypoArena上有所提升,而其他系统包括一个顶级模型出现倒退。与绝对评分细则评分相比,竞技场评估解决了模型之间更细粒度差异,聚合排名与人类专家和独立评判者高度一致。这些结果支持将PHD视为评估大语言模型在不给出最终结论时如何制定调查方向的一个独特目标。我们的代码和数据可在指定网址公开获取。

英文摘要

Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured. We introduce Prospective Hypothesis Discovery (PHD), which asks models to autonomously construct grounded, discriminative, and testable hypothesis spaces from inconclusive evidence, including anomalous observations and fragmented records, to guide subsequent investigation. To evaluate this capability, we introduce HypoArena, comprising HypoData, a benchmark of 988 cases across six scientific and analytical domains, and HypoEval, an evaluation framework for open-ended hypothesis sets. To construct HypoData at scale, we propose Retrospective Context Regression, a Forge--Audit pipeline that reconstructs pre-conclusion contexts from completed expert documents by removing explicit conclusions, target hypotheses, and retrospective causal attributions while preserving the factual substrate. Because PHD admits multiple valid outputs, HypoEval combines bidirectional pairwise judgments with Bradley--Terry--Davidson aggregation for ranking and six-dimensional rubric scoring for diagnosis. Experiments on 15 frontier LLMs reveal clear capability stratification and model-dependent effects of structured analytical skills, with gains for several lower-performing models on HypoArena but regressions for other systems, including a top-performing model. Compared with absolute rubric scoring, arena evaluation resolves finer-grained differences among models, with aggregated rankings showing strong agreement with human experts and an independent judge. Together, these results support treating PHD as a distinct target for evaluating how LLMs formulate investigative directions when final conclusions are withheld. Our code and data are publicly available at github.com/SKYLENAGE-AI/HypoArena and github.com/SKYLENAGE-AI/HypoArena.

URL PDF HTML 收藏
2607.15593 2026-07-20 cs.DC cs.AI cs.NI 新提交

Scalable LLM Agent Tool Access in the Cloud

云中可扩展的大语言模型智能体工具访问

Mingxin Li, Enge Song, Yueshang Zuo, Xiaodong Liu, Rong Wen, Qiang Fu, Gianni Antichi, Jian He, Jing Tie, Zhou Shao, Xiaobo Xue, Xiong Xiao, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Zihao Fan, Haonan Li, Tian Pan, Xiaomin Wu, Yang Song, Xing Li, Biao Lyu, Meng Li, Haipeng Dai, Guihai Chen, Shunmin Zhu

机构 * Nanjing University(南京大学) Alibaba Cloud(阿里云) Fudan University(复旦大学) RMIT University(皇家墨尔本理工大学) Politecnico di Milano(米兰理工学院) Zhejiang University(浙江大学)

AI总结 研究大语言模型智能体在云环境下工具访问问题,提出云规模网关系统,通过打破直接连接模型等整合多种功能,实现高召回率、扩展工具访问数量、提高准确性并减少时间和令牌使用量,还分享了部署经验。

详情
AI中文摘要

大语言模型智能体越来越依赖工具调用与外部系统交互,模型上下文协议(MCP)成为事实上的接口。但在云规模下运行MCP困难重重。工具提供方存在遗留服务难通过MCP直接调用及协议开发带来兼容性成本问题。智能体方面,可访问工具数量受限于大语言模型上下文窗口和推理开销。本文提出云规模的网关系统,打破数据平面直接连接模型,卸载遗留服务集成,整合多种功能。混合检索召回率达98%,将智能体工具访问扩展到3000多个,提高工具选择准确性,减少工具选择时间和令牌使用量,每调用开销低,扩展时稳定。最后分享了生产中部署网关系统的经验教训。

英文摘要

LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de facto interface. Operating MCP at cloud scale, however, becomes difficult. On the tool provider side, legacy services are not directly callable through MCP; the rapid protocol development also creates ongoing compatibility cost. On the agent side, the number of accessible tool is limited by the LLM context window and inference overhead; mounting a large tool set increases token usage and inference latency and can reduce task success rate. Moreover, for stateful MCP backends with multiple replicas, preserving session affinity increases client-side complexity. We present a cloud-scale gateway system for MCP service. It breaks the direct-connect model on the data plane and offloads legacy service integration, consolidating incompatible MCP variants, access control, tool recommendation, and session-aware routing to the gateway. Hybrid retrieval sustains 98% Top-15 recall; it scales agent tool access to 3,000+ with high tool selection accuracy, and reduces tool selection time by $8.9\times$ and token usage by $23.8\times$, with low per-call overhead, stable under scale-out. Finally, we share the lessons learned from deploying the gateway system in production.

URL PDF HTML 收藏
2607.15038 2026-07-20 cs.CV 版本更新

Video = World + Event Stream

视频 = 世界 + 事件流

Lianghua Huang, Zhi-Fan Wu, Yupeng Shi, Wei Wang, Mengyang Feng, Cheng Yu, Chen Liang, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuxiang Bao, Yuzheng Wang, Zhiwei Lin, Zoubin Bi

机构 * Alibaba Group(阿里巴巴集团)

AI总结 研究提出Wan-Streamer v0.3,将视频视为世界加事件流,据此有通用预训练任务,应用于实时全双工视听交互,保留v0.2运行点,为实时下游任务提供能力。

Comments website: https://wan-streamer.com/v0.3/

详情
AI中文摘要

我们展示了Wan-Streamer v0.3,它在单一组织视角下重塑了原生流交互模型:视频是一个世界加一个事件流。世界是视频展开的持久上下文,事件流是世界中随时间变化的一切。这产生了一个针对大量真实视频的通用预训练任务。我们将其应用于实时全双工视听交互。该模型保留了v0.2的运行点,如视频分辨率、帧率等参数。

英文摘要

We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject behavior, speech, and other sounds. This yields a general-purpose pretraining task over large amounts of real video: given a world and incoming input, predict how the world moves, changes, and responds in real time. The resulting competence can be specialized to a broad family of real-time downstream tasks. We instantiate it on real-time full-duplex audio-visual interaction, where the event stream is the agent's speech together with free-form behavior. Functionally, the model's multimodal understanding process is vision-language-action-like: it maps multimodal user input to language-form speech and behavior actions. Wan-Streamer v0.3 preserves the v0.2 operating point: 640x368 video at 25 FPS, a 160 ms streaming unit, approximately 200 ms model-side response latency, and approximately 550 ms total interaction latency under a 350 ms bidirectional network budget.

URL PDF HTML 收藏
2607.09581 2026-07-20 cs.CV cs.SD 版本更新

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

万舞者:一种用于分钟级连贯音乐到舞蹈生成的分层框架

Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Ruoshi Zhang, Yi Lu, Gang Cheng, Bang Zhang

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)

AI总结 针对从音乐生成舞蹈视频的挑战,提出万舞者分层框架,解耦过程为全局关键帧规划和局部时间细化,利用音乐上下文确保连贯,通过动态帧率自适应等创新,突破时长限制,在多舞蹈类型上表现出色,达新的技术水平。

Comments project: https://humanaigc.github.io/wan-dancer-project/, code: https://github.com/Wan-Video/Wan-Dancer, modelscope: https://www.modelscope.cn/models/Wan-AI/Wan-Dancer-14B, huggingface: https://huggingface.co/Wan-AI/Wan2.2-Animate-14B

详情
AI中文摘要

直接从音乐生成长时间、高清且节奏同步的舞蹈视频仍然是一项重大挑战,主要是由于当前扩散模型的时间限制,通常在超过20秒时就会失败。现有方法存在时间漂移、身份不一致和重复运动模式等问题。为此,我们提出了一种用于分钟级连贯音乐到舞蹈生成的新型分层框架。该方法将过程解耦为全局关键帧规划和局部时间细化,利用全轨道音乐上下文确保长程连贯性。关键创新包括通过时间映射的RoPE嵌入进行动态帧率自适应以实现精确对齐、基于光流的损失函数增强运动连续性以及运动速度控制以在快速运动中保留高保真细节。大量实验表明,我们的框架突破了传统时长限制,生成稳定的720p/30fps、超过一分钟的视频,具有卓越的时间稳定性。此外,该模型在五种不同舞蹈类型上表现出强大的通用性,以音频和文本提示为条件,在连贯的长格式舞蹈视频合成方面建立了新的技术水平。

英文摘要

Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether they rely on intermediate 3D skeletons or on end-to-end video synthesis, suffer from temporal drift, identity inconsistency, and repetitive motion patterns when extended to longer horizons. To address these limitations, we propose a novel hierarchical framework for minute-scale coherent music-to-dance generation. Our method decouples the process into global keyframe planning and local temporal refinement, leveraging full-track musical context to ensure long-range coherence. Key innovations include dynamic frame rate adaptation via time-mapped RoPE embeddings for precise alignment, an optical-flow-based loss function to enhance motion continuity, and motion-speed control to preserve high-fidelity details during rapid movements. Extensive experiments demonstrate that our framework surpasses the conventional duration barrier, generating stable, 720p/30fps videos exceeding one minute with superior temporal stability. Furthermore, the model exhibits robust versatility across five distinct dance genres, conditioned on both audio and textual prompts, establishing a new state-of-the-art in coherent, long-form dance video synthesis.

URL PDF HTML 收藏
2509.18127 2026-07-20 cs.LG cs.AI cs.CL

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework

Safe-SAIL: 通过稀疏自编码解释框架构建大语言模型的细粒度安全景观

Jiaqi Weng, Han Zheng, Hanyu Zhang, Ej Zhou, Qinqin He, Jialing Tao, Hui Xue, Zhixuan Chu, Xiting Wang

机构 * Alibaba Group(阿里巴巴集团) The State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江大学区块链与数据安全国家重点实验室) Language Technology Lab, University of Cambridge(剑桥大学语言技术实验室) Renmin University of China(中国人民大学)

AI总结 本文提出Safe-SAIL框架,通过稀疏自编码解释方法高效识别安全领域特征,减少解释成本55%,并系统评估1758个安全相关特征,揭示风险特征识别和安全关键实体编码机制。

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 18916-18935 (2026)

详情
AI中文摘要

稀疏自编码(SAEs)通过将纠缠的模型激活分解为单语义特征,推动可解释性研究。然而,SAEs在何种情况下能为低频概念领域生成最细粒度的潜在特征仍不清楚。本文提出Safe-SAIL框架,旨在安全关键领域解释SAE特征,以提升大语言模型的机理理解。Safe-SAIL引入预解释评估指标,高效识别具有强安全领域可解释性的SAEs,并通过段级模拟策略将解释成本降低55%。基于Safe-SAIL,我们训练了涵盖四个领域(色情、政治、暴力和恐怖)的1758个安全相关特征的综合SAE集合,提供可读解释和系统评估。利用此资源,我们进行了实证分析,探讨了Safe-SAIL在风险特征识别中的有效性,以及安全关键实体和概念在模型层间的编码机制。所有模型、解释和工具均在开源工具包和配套产品中公开发布。

英文摘要

Sparse autoencoders (SAEs) enable interpretability research by decomposing entangled model activations into monosemantic features. However, under what circumstances SAEs derive most fine-grained latent features for safety, a low-frequency concept domain, remains unexplored. Two key challenges exist: identifying SAEs with the greatest potential for generating safety domain-specific features, and the prohibitively high cost of detailed feature explanation. In this paper, we propose Safe-SAIL, a unified framework for interpreting SAE features in safety-critical domains to advance mechanistic understanding of large language models. Safe-SAIL introduces a pre-explanation evaluation metric to efficiently identify SAEs with strong safety domain-specific interpretability, and reduces interpretation cost by 55% through a segment-level simulation strategy. Building on Safe-SAIL, we train a comprehensive suite of SAEs with human-readable explanations and systematic evaluations for 1,758 safety-related features spanning four domains: pornography, politics, violence, and terror. Using this resource, we conduct empirical analyses and provide insights on the effectiveness of Safe-SAIL for risk feature identification and how safety-critical entities and concepts are encoded across model layers. All models, explanations, and tools are publicly released in our open-source toolkit and companion product.

URL PDF HTML 收藏
2510.05750 2026-07-20 cs.LG cs.AI 版本更新

Are Heterogeneous Graph Neural Networks Truly Effective for Node Classification? A Causal Perspective

异构图神经网络在节点分类中真的有效吗?因果视角

Xiao Yang, Xuejiao Zhao, Zhiqi Shen

机构 * College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) Joint NTU-UBC Research Centre of Excellence in Active Living for the Elderly (LILY), Nanyang Technological University(老年人积极生活卓越研究中心(LILY),南洋理工大学) Alibaba-NTU Singapore Joint Research Institute (ANGEL), Nanyang Technological University(阿里巴巴-南洋理工大学新加坡联合研究机构(ANGEL),南洋理工大学)

AI总结 从模型架构和异构信息角度研究异构图神经网络用于节点分类,通过复现实验和因果中介分析框架,发现模型架构和复杂度对性能无因果效应,异构信息通过增加同质性等产生正向因果效应。

详情
AI中文摘要

图神经网络(GNNs)在节点分类方面取得了显著成功。在此基础上,异构图神经网络(HGNNs)整合关系类型以及节点和边的语义以利用异构信息。HGNNs的因果分析发展迅速,旨在区分真实因果效应与虚假相关性。然而,HGNNs在节点分类上是否本质有效仍未得到充分研究,多数研究只是隐含假设而非证实其有效性。在这项工作中,我们从模型架构和异构信息两个角度研究HGNNs用于节点分类。我们在21个数据集和20个基线模型上进行了系统的复现,并进行了全面的超参数重新调整。为进一步厘清性能提升的来源,我们开发了一个因果中介分析框架,将异构关系信息的引入视为处理因素,候选结构属性视为中介变量,节点分类性能视为结果变量。该框架首先根据处理因素引起的变化及其与性能提升的关联筛选候选中介变量,然后将总效应分解为中介效应和直接效应。我们的结果得出两个结论。第一,模型架构和复杂度对节点分类性能没有因果效应。第二,异构信息主要通过增加同质性和局部 - 全局分布差异产生正向因果效应,这使得节点类别更具可区分性。实现代码可在该https网址公开获取。

英文摘要

Graph neural networks (GNNs) have achieved remarkable success in node classification. Building on this progress, heterogeneous graph neural networks (HGNNs) integrate relation types and node and edge semantics to leverage heterogeneous information. Causal analysis for HGNNs is advancing rapidly, aiming to separate genuine causal effects from spurious correlations. However, whether HGNNs are intrinsically effective for node classification remains underexamined, and most studies implicitly assume rather than establish this effectiveness. In this work, we examine HGNNs for node classification from two perspectives: model architecture and heterogeneous information. We conduct a systematic reproduction across 21 datasets and 20 baselines, complemented by comprehensive hyperparameter retuning. To further disentangle the source of performance gains, we develop a causal mediation analysis framework that treats the introduction of heterogeneous relation information as the treatment, candidate structural properties as mediators, and node classification performance as the outcome. This framework first screens candidate mediators according to their treatment-induced changes and their associations with performance improvement, and then decomposes the total effect into mediated and direct effects. Our results lead to two conclusions. First, model architecture and complexity have no causal effect on node classification performance. Second, heterogeneous information exerts a positive causal effect primarily through increasing homophily and local-global distribution discrepancy, which makes node classes more distinguishable. The implementation is publicly available at https://github.com/YXNTU/CausalHGNN.

URL PDF HTML 收藏
2607.14749 2026-07-17 eess.AS cs.CV 新提交

WanSong v1.0 Technical Report

万松v1.0技术报告

Binghui Chen, Pandeng Li, Yu Liu, Jingren Zhou

机构 * Alibaba Group(阿里巴巴集团)

AI总结 针对音乐生成难题,提出万松方法,它是基于扩散的模型,可直接生成高保真多语言长歌曲并输出双声道,通过步长蒸馏加快推理,为微调定制提供途径,助力下游编辑任务。

Comments Wan Team

详情
AI中文摘要

音乐生成基础模型近来备受行业关注。但要实现高效生成、高保真长音频并支持可控性仍具挑战。为此提出万松,一种用于长格式商业级歌曲生成的简单却强大的方法。它是纯基于扩散的模型,能直接生成长达5分钟的高保真多语言歌曲,单次运行输出双声道。其扩散框架通过步长蒸馏实现更快推理,还为微调与定制提供有效途径以支持下游编辑任务。

英文摘要

Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present \textbf{WanSong}, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\eg, AR followed by diffusion), \textbf{WanSong} is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.

URL PDF HTML 收藏
2607.14541 2026-07-17 cs.AI 新提交

Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent

大语言模型生成的GPU内核能否用于生产?基于追踪的基准测试与优化代理

Lingyun Yang, Yuxiao Wang, Shenghao Liang, Linfeng Yang, Daocheng Ying, Chunbo You, Rui Zhang, Luping Wang, Yinghao Yu, Guodong Yang, Liping Zhang

机构 * Alibaba Group(阿里巴巴集团)

AI总结 研究大语言模型生成GPU内核用于生产的可行性,提出Atrex-Bench基准测试,发现现有模型表现不佳。为此发布Atrex-Kernel-Agent优化代理,结合多种技术,经案例研究能将回退转换为匹配或超越生产基线的内核。

Comments Both artifacts are released as open source: Atrex-Bench (https://github.com/alibaba/atrex-bench) and Atrex-Kernel-Agent (https://github.com/alibaba/atrex-kernel-agent)

详情
AI中文摘要

现有的GPU内核生成基准测试的问题来源于合成或精心策划的来源,与实际部署的工作负载不同。我们提出了Atrex-Bench基准测试,其30个运算符和440种形状直接从计算受限、内存丰富的GPU的全集群生产推理追踪中采样。每个问题都有一个重要性权重,通过应用卡小时加权,并针对其运行的服务阶段单独计算,还有每个问题的屋顶线上限。评估六个前沿编码代理表明,即使是最好的原始模型在生产运算符上也只能达到硬件屋顶线的约10%。为了缩小差距,我们共同发布了Atrex-Kernel-Agent(AKA),它结合了迭代测量-修正搜索、用于避免搜索上下文停滞的优化随机失活,以及分层的GPU优化知识库。在一个受控案例研究中,该代理将零FlyDSL回退转换为与手工调整的生产基线匹配或超过的实际内核。

英文摘要

Existing GPU kernel generation benchmarks draw problems from synthetic or curated sources that diverge from deployed workloads. We present Atrex-Bench, a benchmark whose 30 operators and 440 shapes are sampled directly from full-cluster production inference traces of compute-limited, memory-rich GPUs. Each problem carries an importance weight derived from its share of observed GPU time, weighted by application card-hours and computed separately for the serving phases in which it runs, together with a per-problem roofline ceiling, so the aggregate score emphasizes the kernels that consume the most serving time. Evaluating six frontier coding agents on Atrex-Bench shows that even the best vanilla model reaches only ${\sim}10\%$ of the hardware roofline on production operators; and correctness alone overstates capability, since much of the apparent pass rate comes from PyTorch fallbacks rather than kernels the model wrote. To close this gap, we co-release Atrex-Kernel-Agent (AKA), a profile-driven kernel-optimization agent that combines iterative measure-revise search, optimization dropout for escaping stalled search contexts, and a layered GPU-optimization knowledge base (298 reference-kernel files and 244 optimization-knowledge documents, plus external upstream reference projects for API/ISA lookup). In a controlled case study, the agent converts zero-FlyDSL fallbacks into real kernels that match or exceed hand-tuned production baselines.

URL PDF HTML 收藏
2605.28732 2026-07-17 cs.CL cs.AI cs.LG 版本更新

MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems

MemTrace:大型语言模型记忆系统中的错误追踪与归因

Xinle Deng, Ruobin Zhong, Hujin Peng, Xiaoben Lu, Yanzhe Wu, Guang Li, Buqiang Xu, Yunzhi Yao, Jizhan Fang, Haoliang Cao, Junjie Guo, Yuan Yuan, Ziqing Ma, Yuanqiang Yu, Rui Hu, Baohua Dong, Hangcheng Zhu, Ningyu Zhang

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团)

AI总结 提出MemTrace框架,通过构建可执行的记忆演化图实现细粒度错误追踪,并利用自动归因方法定位根因,进而优化提示词提升下游任务性能。

Comments Ongoing work

详情
AI中文摘要

记忆对于使大型语言模型支持长程推理至关重要,但现有的记忆系统仍然不可靠且难以调试。追踪记忆的动态演化对于理解信息如何随时间合成、传播或损坏至关重要。在这项工作中,我们研究了LLM记忆系统中错误追踪与归因的新问题。我们提出了一种新颖的框架,将记忆流水线转换为可执行的记忆演化图,从而实现对操作信息流的细粒度追踪。然后,我们构建了MemTraceBench,一个从代表性记忆系统(如Long-Context、RAG、Mem0和EverMemOS)收集的基准,以系统地研究记忆故障模式。我们进一步引入了一种自动归因方法,该方法迭代地追踪操作子图以定位任何失败案例的根本原因。我们的分析表明,记忆故障是系统性的,源于操作层面的问题,如信息丢失和检索错位。关键的是,我们利用这些细粒度的归因信号来指导下游提示优化,建立了一个自动纠正故障并提升最终任务性能高达7.62%的闭环系统。代码将在https://github.com/zjunlp/MemTrace发布。

英文摘要

Memory is essential for enabling large language models to support long-horizon reasoning, yet existing memory systems remain unreliable and difficult to debug. Tracing memory's dynamic evolution is crucial to understand how information is synthesized, propagated, or corrupted over time. In this work, we study the new problem of error tracing and attribution in LLM memory systems. We propose a novel framework that transforms memory pipelines into executable memory evolution graphs, enabling fine-grained tracing of operational information flow. We then construct MemTraceBench, a benchmark collected from representative memory systems such as Long-Context, RAG, Mem0, and EverMemOS, to systematically study memory failure modes. We further introduce an automatic attribution method that iteratively traces operation subgraphs to pinpoint the root cause of any failed case. Our analysis reveals that memory failures are systematic, stemming from operation-level issues like information loss and retrieval misalignment. Crucially, we leverage these fine-grained attribution signals to guide downstream prompt optimization, establishing a closed-loop system that automatically corrects faults and boosts end-task performance by up to 7.62%. Code will be released at https://github.com/zjunlp/MemTrace.

URL PDF HTML 收藏
2605.30060 2026-07-17 cs.CV 版本更新

Towards Consistent Video Geometry Estimation

Towards Consistent Video Geometry Estimation

Zhu Yu, Jingnan Gao, Runmin Zhang, Lingteng Qiu, Zhengyi Zhao, Rui Peng, Yichao Yan, Kejie Qiu, Siyu Zhu, Zilong Dong, Si-Yuan Cao, Hui-Liang Shen

机构 * Zhejiang University(浙江大学) Tongyi Lab, Alibaba Group(阿里云实验室) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学)

AI总结 提出ViGeo,一种基于纯Transformer架构的前馈基础模型,通过动态分块注意力机制和基于补全的数据精炼框架,实现视频序列中空间密集且时间一致的几何(深度、法线、点图)估计,在在线、离线及长视频任务中达到最先进性能。

Comments Project webpage: https://pkqbajng.github.io/ViGeo/

详情
AI中文摘要

本文提出了ViGeo,一种前馈基础模型,用于从视频序列中恢复空间密集且时间一致的几何信息。ViGeo基于纯Transformer架构,没有针对特定任务的架构修改,支持在统一模型中进行流式、全序列和长视频推理。关键设计是动态分块注意力,该机制在训练期间使模型同时暴露于双向和因果时间上下文,并允许其在测试时无需重新训练即可调整注意力模式。为了提高监督质量,我们进一步引入了一种基于补全的数据精炼框架。该框架训练了一个视频深度补全教师模型,该模型以稀疏且有噪声的标注为条件,利用视频/多视图上下文生成密集、时间一致且几何可靠的训练目标。除了深度和点图,ViGeo还在同一框架内预测表面法线。仅使用公共数据集训练,ViGeo在在线、离线和长视频深度估计、表面法线估计以及视频点图估计中均达到了最先进性能。

英文摘要

This work presents ViGeo, a feed-forward foundation model for recovering spatially dense and temporally consistent geometry from video sequences. Built upon a plain transformer architecture without task-specific architectural modifications, ViGeo supports streaming, full-sequence, and long-video inference within a unified model. The key design is dynamic chunking attention, which exposes the model to both bidirectional and causal temporal contexts during training and allows it to adapt its attention pattern at test time without retraining. To improve supervision quality, we further introduce a completion-based data refinement framework. This framework trains a video depth completion teacher that conditions on sparse and noisy annotations and exploits video/multi-view context to produce dense, temporally coherent, and geometrically reliable training targets. Beyond depth and point maps, ViGeo also predicts surface normals within the same framework. Trained solely on public datasets, ViGeo achieves state-of-the-art performance across online, offline, and long-video depth estimation, surface normal estimation, and video point map estimation.

URL PDF HTML 收藏
2607.13941 2026-07-16 cs.CV 新提交

Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment

峰值-结尾网络:一种受峰值-结尾规则启发的通用视频美学评估框架

Geng Li, Haiwen Li, Rui Chen, Jing Tang, Lei Sun, Xiangxiang Chu

机构 * Alibaba Group(阿里巴巴集团) Beijing University of Posts and Telecommunications(北京邮电大学)

AI总结 研究视频美学评估问题,提出受峰值-结尾规则启发的Peak-End-Net框架,通过引入预训练IAA头部、设计美学节奏编码器和动态门控融合机制,基于冻结ViT实现,在实验中取得最优性能。

Comments Accepted to ACM MM 2026, Code: https://github.com/AMAP-ML/Peak-End-Net

详情
AI中文摘要

视频美学评估(VAA)旨在预测视频的美学吸引力,但与其他视觉评估任务相比,其探索较少。其进展受到大规模基准稀缺以及美学判断内在主观性的阻碍。本文从心理学角度重新审视VAA,提出了受峰值-结尾规则启发的轻量级且可解释的框架Peak-End-Net。通过引入预训练的图像美学评估(IAA)头部来生成逐帧美学先验,设计美学节奏编码器以及动态门控融合机制,该方法基于冻结的视觉Transformer(ViT),参数少且可扩展。在两个现有VAA基准上的大量实验表明其达到了当前最优性能。

英文摘要

Video aesthetic assessment (VAA) aims to predict how aesthetically pleasing a video is, yet remains far less explored than other visual assessment tasks. Its progress is hindered not only by the scarcity of large-scale benchmarks, but also by the intrinsic subjectivity of aesthetic judgment, which is shaped by human perception. In this paper, we revisit VAA from a psychological perspective and propose \textit{Peak-End-Net}, a lightweight and interpretable framework inspired by the \textit{peak-end rule}, which suggests that people tend to judge a temporal experience mainly according to its salient moments and the ending. Building on this intuition, we first transfer knowledge from image aesthetic assessment (IAA) to VAA by introducing a pretrained IAA head to produce frame-wise aesthetic priors, which serve as surrogate signals for identifying aesthetically salient moments and guiding \textit{peak-end rule}-based temporal aggregation. To further capture how a video evolves aesthetically over time, we design an aesthetic rhythm encoder that models temporal progression beyond isolated moments. Additionally, we refine the overall assessment through a dynamic gated fusion mechanism to improve robustness under distribution shift. Our method is built on a frozen vision transformer (ViT) and requires only a small number of trainable parameters, making it scalable and parameter-efficient. Extensive experiments on two existing VAA benchmarks, including in-domain evaluation on VADB and cross-domain testing on DIVIDE-3K, demonstrate that our approach achieves state-of-the-art performance, affirming the value of psychologically grounded modeling for VAA. Our code and models are available at https://github.com/AMAP-ML/Peak-End-Net.

URL PDF HTML 收藏
2607.13639 2026-07-16 cs.CV cs.AI 新提交

OvisOCR2 Technical Report

OvisOCR2技术报告

Shiyin Lu, Yinglun Li, Yu Xia, Yuhui Chen, An-Yang Ji, Jun-Peng Jiang, Qing-Guo Chen, Jianshan Zhao, En Lin, Haijun Li, Cheng Qin, Zhao Xu, Weihua Luo

机构 * Alibaba Group(阿里巴巴集团)

AI总结 介绍拥有8亿参数的OvisOCR2文档解析模型,通过构建数据引擎,采用监督微调、强化学习、策略蒸馏和模型融合等方法训练,在多个基准测试中取得优异成绩,展现出良好的泛化性和鲁棒性。

详情
AI中文摘要

我们介绍了OvisOCR2,一个拥有8亿参数的文档解析模型。它被设计为端到端解析器,给定文档页面图像,能按自然阅读顺序生成Markdown表示,涵盖文本、公式、表格和视觉区域。我们构建了数据引擎,结合了经过筛选的真实文档注释与合成页面。训练方法包括监督微调、在一个拥有40亿参数分支上进行多组件奖励设计的强化学习、策略蒸馏到8亿参数模型以及模型融合。在OmniDocBench v1.6上,OvisOCR2取得了96.58的最优综合得分,在PureDocBench上也取得了75.06的最高Avg3得分。在内部基准测试中,OvisOCR2在比较方法中获得了最佳整体性能,证明了其泛化性和鲁棒性。

英文摘要

We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generates a Markdown representation in natural reading order, covering text, formulas, tables, and visual regions. We build a data engine that combines filtered real-document annotations with synthetic pages whose rendered images and Markdown targets are derived from the same HTML source. The training recipe includes supervised fine-tuning, reinforcement learning on a 4B branch with a multi-component reward design, on-policy distillation into the 0.8B model, and model fusion. On OmniDocBench v1.6, OvisOCR2 achieves a state-of-the-art overall score of 96.58, placing an end-to-end model at the top of this leaderboard previously dominated by pipeline methods and highlighting the potential of end-to-end document parsing. On PureDocBench, OvisOCR2 also achieves the highest Avg3 score of 75.06. Beyond these two public benchmarks, we evaluate OvisOCR2 on an in-house benchmark designed to cover a broader set of long-tail and challenging scenarios. OvisOCR2 obtains the best overall performance among the compared methods, providing further evidence of its generalization and robustness. OvisOCR2 is available at https://huggingface.co/ATH-MaaS/OvisOCR2.

URL PDF HTML 收藏
2607.13017 2026-07-15 cs.RO cs.CV 新提交

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

FlowWAM:光流作为世界动作模型的统一动作表示

Yixiang Chen, Peiyan Li, Yuan Xu, Qisen Ma, Jiabing Yang, Kai Wang, Jianhua Yang, Dong An, He Guan, Gaoteng Liu, Jianlou Si, Jun Huang, Jing Liu, Nianfeng Liu, Yan Huang, Liang Wang

机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) FiveAges(无) MBZUAI(无) Alibaba Group(阿里巴巴集团)

AI总结 研究针对世界动作模型控制中动作表示难题,提出FlowWAM双流扩散框架,以光流为统一动作表示。该框架可实现WAMs两种模式,能利用无动作标签视频预训练,实验表明在操纵和世界建模任务中表现优于基线。

详情
AI中文摘要

世界动作模型(WAMs)可利用预训练视频生成器进行世界建模和动作预测。但直接用于控制面临挑战:如何以合适形式表示动作,既与预训练视频生成器匹配,又携带足够运动线索用于精确控制。现有数值动作和视觉动作表示均有不足。我们提出FlowWAM,一种双流扩散框架,采用光流作为统一的、视频原生的动作表示。通过在共享预训练视频生成器中联合建模,FlowWAM可实现WAMs的两种模式。在策略模式下用于动作预测,在世界模型模式下用目标流序列指导未来视频生成。此外,光流可从无动作标签的原始视频中轻松提取,能利用大规模无动作标签视频数据集进行预训练。实验表明,基于光流的动作表示在两种模式下均有提升。在RoboTwin操纵任务中,在Clean设置下成功率达92.94%,在Random设置下为92.14%,优于VLA和WAM基线。在WorldArena世界建模任务中,实现最佳总体EWMScore(63.71),轨迹精度相对提高18.4%。更多结果可在项目网站查看。

英文摘要

World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control. Existing numerical actions fail to satisfy the former, and prior visual action representations overlook the temporal motion structure across frames. We address this issue with FlowWAM, a dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Flow videos share the same format as RGB videos and encode rich per-pixel displacement. By jointly modeling them within a shared pretrained video generator, FlowWAM can naturally implement two modes of WAMs. In policy mode, FlowWAM generates flow for action prediction, while in world-model mode, it uses target flow sequences to guide future video generation. Moreover, since flow can be easily extracted from raw videos without action labels, FlowWAM can leverage large-scale action-unlabeled video datasets for pretraining. We empirically find that our flow-based action representation delivers gains across both modes. On RoboTwin manipulation, FlowWAM raises the success rate to 92.94% on the Clean setting and 92.14% on Random, outperforming both VLA and WAM baselines. On WorldArena world modeling, it achieves the best overall EWMScore (63.71) with an 18.4% relative improvement in trajectory accuracy. More results can be found on our project website: https://flow-wam.github.io .

URL PDF HTML 收藏
2607.12680 2026-07-15 cs.CV 新提交

ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning

ReflectVLN:通过反思推理训练视觉语言导航智能体

Jiahang Wang, Yirong Yang, Yanqing Zhu, Minghua Luo, Shichao Xie, Fei Liu, Mu Xu

机构 * Amap, Alibaba Group(高德软件有限公司,阿里巴巴集团) Beihang University(北京航空航天大学)

AI总结 研究针对现有视觉语言导航方法缺乏闭环机制问题,提出ReflectVLN框架,通过双向交互智能体决策,引入行动思维链训练方案,实验证明该框架在有限数据下提升成功率和路径效率,且具良好训练成本与可解释性。

详情
AI中文摘要

现有视觉语言导航方法常将视觉语言模型(VLM)与航点解码器结合生成多步行动计划,但缺乏明确闭环机制来跟踪语义进展、诊断执行失败及从长期导航中的错误积累中恢复。为填补这一空白,我们提出ReflectVLN,一个通过双向交互意图和执行智能体组织决策的智能体视觉语言导航框架。意图智能体执行子任务分解和反思,生成可执行的子任务描述作为纠正计划。执行智能体依据这些描述在当前观察下将其转化为短期行动,同时监测子目标进展并检测偏离行为。关键的是,ReflectVLN实现了闭环双向通信。为鼓励具有可解释中间推理的时间连贯决策,我们引入行动思维链(Action-CoT),一种用于行动生成的路径条件双查询训练方案。在标准视觉语言导航基准测试中的实验表明,ReflectVLN在有限数据预算下提高了成功率和路径效率,具有良好的训练成本,推理时高级意图调用更少,同时提供可解释的中间决策用于分析和协作。

英文摘要

Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution failures, and recovering from error accumulation in long-horizon navigation. To address this gap, we propose ReflectVLN, an agentic VLN framework that organizes decision-making through bidirectionally interactive intention and execution agents. The intention agent performs subtask decomposition and reflection, generating executable subtask descriptions as corrective plans. Conditioned on these descriptions, the execution agent grounds them into short-horizon actions under current observations while monitoring sub-goal progress and detecting off-track behavior. Crucially, ReflectVLN enables closed-loop bidirectional communication: the execution agent emits progress and deviation signals to trigger reflection and subtask updates on demand, and the intention agent returns structured guidance that reconditions subsequent actions for recovery. To encourage temporally coherent decisions with interpretable intermediate rationales, we introduce Action Chain-of-Thought (Action-CoT), a path-conditioned dual-query training scheme for action generation. Experiments on standard VLN benchmarks show that ReflectVLN improves success rates and path efficiency under a constrained data budget, with favorable training cost and fewer high-level intention calls at inference time, while providing interpretable intermediate decisions for analysis and collaboration. Code is available at: https://github.com/AIprogrammer/ReflectVLN

URL PDF HTML 收藏
2607.12592 2026-07-15 cs.CV 新提交

WanToFight: Real-Time Generative Game Engine for Multi-Player Combat Interaction

WanToFight:用于多人战斗交互的实时生成游戏引擎

Li Hu, Guangyuan Wang, Peng Zhang, Bang Zhang

机构 * Alibaba Tongyi Lab(阿里巴巴通义实验室)

AI总结 WanToFight是用于多人战斗交互的实时生成游戏引擎,基于Wan-1.3B视频扩散变换器构建三个组件,解决了先前引擎未共同处理的问题,能结合多种玩法,在特定配置下维持30FPS,是首个此类集成系统。

Comments Project Page: https://humanaigc.github.io/wantofight/

详情
AI中文摘要

我们展示了WanToFight,这是一个生成式游戏引擎,可根据键盘输入模拟实时两人的《拳皇97》游戏玩法。先前的生成式游戏引擎要么针对单人第一人称设置,要么针对非实时合作场景;多人控制、实时推理、复杂的物理交互和对抗性游戏玩法尚未得到共同解决。WanToFight通过基于Wan-1.3B视频扩散变换器构建的三个组件填补了这一空白:具有块因果注意力和滚动KV缓存的流式自回归生成器;一个视觉基础的玩家关联模块,将每个玩家的键盘信号绑定到一个角色身份;以及一个通过单人到全游戏课程训练的门控、局部因果键盘注入模块。一个经过四步DMD提炼的学生与一个剪枝的VAE解码器在单个NVIDIA RTX 5090上以512x384的分辨率在完整比赛期间维持30FPS。据我们所知,WanToFight是第一个在一个系统中结合多人控制、实时推理、复杂物理交互和对抗性游戏玩法的生成式游戏引擎。

英文摘要

We present WanToFight, a generative game engine that simulates real-time, two-player The King of Fighters '97 (KOF~'97) gameplay from keyboard input. Prior generative game engines target either single-player first-person settings or non-real-time cooperative scenarios; multi-player control, real-time inference, complex physical interaction, and adversarial gameplay have not been jointly addressed. WanToFight closes this gap with three components built on the Wan-1.3B video diffusion transformer: a streaming autoregressive generator with block-causal attention and a rolling KV cache; a visually grounded Player Association module that binds each player's keyboard signal to a character identity; and a gated, locally causal keyboard injection module trained with a single-player-to-full-gameplay curriculum. A four-step DMD-distilled student paired with a pruned VAE decoder sustains 30FPS at 512x384 on a single NVIDIA RTX 5090 over the duration of a complete match. To our knowledge, WanToFight is the first generative game engine to combine multi-player control, real-time inference, complex physical interaction, and adversarial gameplay in one system.

URL PDF HTML 收藏
2607.11019 2026-07-15 cs.AI 版本更新

QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics

QwenPaw-数据:为自主企业数据分析搭建事实、方法与执行的桥梁

Tianjing Zeng, Yuntao Hong, Zhongjun Ding, Dandan Liu, Yinan Mei, Yunxiang Su, Yiming Wang, Xiaojian Zhang, Jingyu Zhu, Junhao Zhu, Zhuowen Liang, Jiazhen Peng, Lianggui Weng, Zhihao Ding, Kerui Yi, Qifeng Wang, Rong Zhu, Bolin Ding, Liyu Mou, Jingren Zhou

机构 * Alibaba Group(阿里巴巴集团)

AI总结 研究针对企业数据分析在开放环境的特性,提出QwenPaw-Data系统,通过整合异构资产、转化自然语言请求为工作流,架构含三个协作子系统,实验证明该系统提升了数据访问和分析质量,为企业数据智能体提供基础。

详情
AI中文摘要

企业数据分析正成为自主智能体的一个独特前沿领域。与通用交互和软件工程相比,它在开放、模糊且不断演变的环境中运行。这些特性需要一种将语义、方法、执行和演变作为首要系统关注点的数据智能体架构。为此,我们引入了QwenPaw-Data,一个为企业智能数据分析设计的智能体数据系统。它整合来自仓库、仪表盘、文档、交互日志和历史任务的异构资产,将自然语言请求转化为端到端分析工作流。其架构分解为三个协作子系统:DataBridge通过互连的元数据、知识和跟踪图提供可靠语义基础;Skill-Hub将专家分析方法编纂为可复用和可验证技能;Host将这些证据和方法资产转化为可控的、以工件为中心的运行时执行。实验表明它提高了可验证数据访问能力和高级分析质量,为企业数据智能体提供了实用基础。

英文摘要

Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambiguous, and continuously evolving environment. These characteristics call for a data-agent architecture that treats semantics, methodology, execution, and evolution as first-class system concerns. To this end, we introduce QwenPaw-Data, an agentic data system designed for enterprise intelligent data analysis. QwenPaw-Data consolidates heterogeneous assets from warehouses, dashboards, documents, interaction logs, and historical tasks into reusable, governable, and evolvable analysis assets, then turns natural-language requests into end-to-end analytical workflows spanning data understanding, retrieval, analysis, report generation, and decision support. Its architecture decomposes the problem into three collaborative subsystems: DataBridge provides trustworthy semantic grounding through interconnected metadata, knowledge, and trace graphs; Skill-Hub codifies expert analytical methodology into reusable and verifiable skills; and Host materializes these evidence and method assets into controllable, artifact-centric runtime execution. Across these subsystems, semantics, methods, traces, and feedback are continuously deposited back into the system, forming a self-evolving asset flywheel. Experiments on public benchmarks and real-world industrial BI workloads show that QwenPaw-Data improves both verifiable data access capability and higher-level analytical quality, offering a practical foundation for reliable, traceable, and continuously improving enterprise data agents.

URL PDF HTML 收藏
2606.09076 2026-07-15 cs.CV 版本更新

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions

超越标量奖励:将推理内化到分数分布中

Xin Jin, Huanqia Cai, Zhen Li, Zechao Zhan, Dengyang Jiang, Aiming Hao, Yuming Jiang, Xiangpeng Yang, Chunle Guo, Peng Gao, Ming-Ming Cheng, Steven C. H. Hoi

机构 * Alibaba Group(阿里巴巴集团) Nankai University(南开大学)

AI总结 提出Z-Reward框架,通过教师-学生模型将推理型奖励内化为紧凑VLM的分数分布,实现高效且准确的文本到图像优化。

Comments Z-Image Team Technical Report. Project page: https://srameo.github.io/projects/z-reward/

详情
AI中文摘要

奖励模型对于文本到图像的后训练至关重要,但视觉偏好是主观的,更适合表示为评分分布而非确定性标量。现有的标量、评分令牌和成对奖励模型过度压缩了不确定性和细粒度评分差异,而基于推理的生成式奖励提供了更强的判断,但部署成本高且难以用作直接优化信号。我们提出Z-Reward,一种教师-学生奖励建模框架,将推理密集型判断与高效奖励部署解耦。教师是一个大型VLM,使用推理推断符合评分标准的分数分布,并通过组定向分数优化(GDSO)进行训练,该优化结合了来自分布期望的策略梯度奖励以及关于分数分布和分数差距的直接点式和成对监督。学生通过推理内化分数蒸馏(RISD)进行训练,将教师的推理条件分数分布转移到紧凑VLM中,而无需在推理时使用显式推理链。在我们内部标注的评估集上,27B GDSO教师达到了89.6%的人类偏好准确率,优于SFT、RewardDance和GRPO,而9B RISD学生达到了88.6%,优于OPD基线并接近更大的教师。我们进一步表明,Z-Reward可以作为文本到图像优化的可微奖励信号,相对于SFT基线产生了41.3%的净人类偏好改进。

英文摘要

Reward models are central to text-to-image post-training, but visual preference is subjective and better represented as a distribution over rubric scores than as a deterministic scalar. Existing scalar, score-token, and pairwise reward models over-compress uncertainty and fine-grained score differences, while reasoning-based generative rewards provide stronger judgments but are costly to deploy and difficult to use as direct optimization signals. We propose Z-Reward, a teacher-student reward modeling framework that decouples reasoning-heavy judgment from efficient reward deployment. The teacher is a large VLM that uses reasoning to infer rubric-aligned score distributions, and is trained with Group-wise Direct Score Optimization (GDSO), which combines policy-gradient rewards from distribution expectations with direct pointwise and pairwise supervision on score distributions and score gaps. The student is trained with Reasoning-Internalized Score Distillation (RISD), which transfers the teacher's reasoning-conditioned score distribution into a compact VLM without requiring explicit reasoning chains at inference time. On our internally annotated evaluation set, the 27B GDSO teacher reaches 89.6% human preference accuracy, outperforming SFT, RewardDance, and GRPO, while the 9B RISD student reaches 88.6%, outperforming the OPD baseline and closely matching the larger teacher. We further show that Z-Reward can serve as a differentiable reward signal for text-to-image optimization, yielding a 41.3% net human-preference improvement over the SFT baseline.

URL PDF HTML 收藏
2606.03363 2026-07-15 cs.CL 版本更新

EntSQL: A Benchmark for Grounding Text-to-SQL in Long-Context Enterprise Knowledge

EntSQL:一个将Text-to-SQL置于长上下文企业知识中的基准

Chengxi Liao, Tao Xu, Zulong Chen, Chuanfei Xu, Yiyan Wang, Xinyun Wang, Yanlong Zhang, Xiaojun Chen, Zhibo Yang, Zeyi Wen

机构 * HKUST (GZ)(香港科技大学(广州)) Alibaba Group(阿里巴巴集团) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳))

AI总结 提出EntSQL基准,通过包含1066个跨五个业务领域的中英文对齐示例,评估LLM在长上下文企业文档中基于私有业务知识生成SQL的能力,最佳系统仅达15.9%准确率。

详情
AI中文摘要

Text-to-SQL使得通过自然语言访问数据库成为可能,最近的LLM显著提升了其能力。现有的基准如Spider、BIRD和Spider~2.0评估了模式泛化、大规模数据库和现实工作流,但很大程度上忽略了SQL生成依赖于私有业务知识(如内部指标、报告惯例和组织规则)的企业场景。我们引入了EntSQL,一个面向企业的Text-to-SQL基准,用于评估在专有业务文档上的长上下文基础。EntSQL包含1066个跨五个业务领域的中英文对齐语义示例,大多数示例需要超越问题和模式的领域知识,并涉及复杂的SQL结构。在英文输入上,当提供长文档时,最佳评估系统仅达到15.9%,突显了在企业知识基础上生成SQL的难度。

英文摘要

Text-to-SQL enables natural language access to databases, and recent LLMs have substantially advanced its capabilities. Existing benchmarks such as Spider, BIRD, and Spider~2.0 evaluate schema generalization, large-scale databases, and realistic workflows, but largely overlook enterprise scenarios where SQL generation depends on private business knowledge, such as internal metrics, reporting conventions, and organizational rules. We introduce EntSQL, an enterprise-oriented Text-to-SQL benchmark for evaluating long-context grounding over proprietary business documents. EntSQL contains 1,066 aligned Chinese-English semantic examples across five business domains, with most examples requiring domain knowledge beyond the question and schema and involving complex SQL structures. On English inputs, the best evaluated system reaches only 15.9\% when long-form documents are provided, highlighting the difficulty of grounding SQL generation in enterprise knowledge.

URL PDF HTML 收藏