arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Beihang University(北京航空航天大学)

至 收录 1216
2607.18016 2026-07-21 cs.RO 新提交

Closing the Loop in Humanoid VLA: Persistent 3D Object Tokens for Verifiable Loco-Manipulation

人形机器人视觉语言动作闭环:用于可验证移动操作的持久3D对象令牌

Peng Ren, Haoyang Ge, Jiang Zhao, Cong Huang, Yukun Shi, Pei Chi, Kai Chen

机构 * BUAA(北京航空航天大学) DeepCybo(深灵机器人) ZGCI(未提及具体中文译名)

AI总结 研究人形机器人视觉语言动作中对象状态差异问题,提出持久对象令牌化方法(POT)并实例化为POT-VLA,通过RGB-D观察维护3D对象记录,转换为对象令牌用于动作生成与验证,实验表明该方法提升了任务成功率。

详情
AI中文摘要

视觉语言动作策略是通用机器人控制的一个有前景的基础,但长期的人形机器人移动操作要求机器人在运动、接触、遮挡和恢复过程中将任务对象视为持久的物理实体。我们将此问题研究为对象状态差异:用于调节全身动作的对象状态可能与用于判断动作是否实现预期物理关系的状态不同。我们提出了“持久对象令牌化”(POT),它从RGB-D观察中维护按角色索引的3D对象记录,并将其转换为用于全身动作专家的对象令牌。实例化为“POT-VLA”时,相同的对象记录可调节动作生成并支持几何谓词检查,产生一个闭环执行系统,其中对象状态既具有可操作性又具有可验证性。在Unitree G1上,POT-VLA将匹配的直接GR00T-N1.7基线在八个现实世界任务族中的成功率从39/80提高到71/80。在外部与Being-0对齐的参考中,POT-VLA在对齐的服务任务上实现了44/50的成功率,而Being-0论文报告的成功率为37/50。最大的收益出现在需要维持3D关系的任务上,这表明持久的以对象为中心的状态是可验证人形机器人视觉语言动作执行的有用抽象。

英文摘要

Vision-language-action policies are a promising foundation for general robot control, but long-horizon humanoid loco-manipulation requires the robot to treat task objects as persistent physical entities across movement, contact, occlusion, and recovery. We study this problem as object-state divergence: the object state used to condition a whole-body action can differ from the state used to decide whether the action achieved the intended physical relation. We propose \emph{Persistent Object Tokenization} (POT), which maintains role-indexed 3D object records from RGB-D observations and converts them into object tokens for a whole-body action expert. Instantiated as \emph{POT-VLA}, the same object records condition action generation and support geometric predicate checks, yielding a closed-loop execution system in which object state is both actionable and verifiable. On a Unitree G1, POT-VLA improves a matched direct GR00T-N1.7 baseline from 39/80 to 71/80 successes over eight real-world task families. In an external Being-0-aligned reference, POT-VLA achieves 44/50 successes on aligned service tasks, compared with the 37/50 success reported by the Being-0 paper. The largest gains occur on tasks requiring maintained 3D relations, suggesting that persistent object-centered state is a useful abstraction for verifiable humanoid VLA execution.

URL PDF HTML 收藏
2607.17768 2026-07-21 cs.CV 新提交

To Blend In, First Decouple: Rethinking Camouflage Image Generation via Context-Decoupled Representations

为了融入,首先解耦:通过上下文解耦表示重新思考伪装图像生成

Wenzhuang Wang, Yifan Zhao, Mingcan Ma, Yunlong Che, Haoran Chen, Ming Liu, Jia Li

机构 * Beihang University(北京航空航天大学) Geely Automobile Research Institute (Ningbo) Co., Ltd(吉利汽车研究院(宁波)有限公司)

AI总结 研究伪装图像生成问题,提出上下文解耦生成范式CamoDreamer,通过设计对比感知上下文桥等方法,将潜在伪装特征解耦为物体和背景控制流,实验证明其显著优于现有方法且设计轻量。

Comments 14 pages, 11 figures, ACMMM 2026

详情
AI中文摘要

伪装图像生成(CIG)专注于生成视觉上隐藏的物体,使其无缝融入背景。现有方法通常遵循背景引导范式或前景引导策略,但仍存在外观差异和背景伪影问题。我们将这些限制归因于跨上下文表示泄漏。为此,我们提出了一种新的上下文解耦生成范式CamoDreamer,旨在隔离上下文条件引导并将潜在伪装特征解耦为协调的物体和背景控制流。具体包括设计对比感知上下文桥、使用上下文解耦同化流、频率自适应上下文混合模块。实验表明CamoDreamer显著优于现有方法且设计相对轻量。

英文摘要

Camouflage image generation (CIG) focuses on generating visually concealed objects that seamlessly blend into their backgrounds. Existing methods typically follow either background-guided paradigms that adapt object appearance via style transfer, or foreground-guided strategies that outpaint surrounding regions conditioned on object features. However, they still suffer from appearance discrepancy and background artifacts. We attribute these limitations to cross-context representation leakage, where object and background cues are entangled in a coupled conditional space, resulting in ambiguous control and degraded camouflage fidelity. To tackle this, we propose a new context-decoupled generative paradigm, termed CamoDreamer, which aims to isolate contextual conditional guidance and explicitly decouple latent camouflage features into coordinated object and background control streams. First, a Contrast-aware Contextual Bridge is designed to model cross-context discrepancies and construct contrast-aware dual conditional guidance. Second, Context-Decoupled Assimilation Streams are employed to separate generative interactions conditioned on the dual guidance, while facilitating background rendering with target-aware cues in the latent space. Finally, a Frequency-Adaptive Contextual Blend module integrates complementary high-frequency textures and low-frequency structures from decoupled features to improve holistic coherence. Extensive experiments demonstrate that CamoDreamer consistently outperforms existing methods with a substantial margin, while maintaining a relatively lightweight design.

URL PDF HTML 收藏
2607.17069 2026-07-21 cs.CV 新提交

AdvSerial: Physical Adversarial Attacks on Infrastructure-mounted Pedestrian Detectors via Semantic Feature Suppression

AdvSerial:通过语义特征抑制对基于基础设施的行人检测器进行物理对抗攻击

Yuanhao Huang, Yilong Ren, Jinlei Wang, Xuesong Bai, Jinchuan Zhang, Haiyang Yu

机构 * School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院) State Key Lab of Intelligent Transportation System(智能交通系统国家重点实验室) Zhongguancun Laboratory(中关村实验室)

AI总结 针对基础设施行人检测器安全漏洞,提出AdvSerial框架,通过2D-3D联合优化、语义特征抑制等方法生成对抗补丁,在多检测器实验中成功率高、可转移性强,揭示故障模式,为相关防御设计提供思路。

详情
AI中文摘要

基于人工智能的视觉感知系统越来越多地应用于基础设施监控中,但其易受物理对抗攻击的安全漏洞对交通基础设施的可靠运行构成直接威胁。本文提出AdvSerial,一个动态的2D-3D联合优化框架,用于在基于基础设施的场景中针对行人检测器生成连续的高角度物理对抗补丁。通过UV映射将边界感知的拼接纹理映射到3D服装上,结合2D数字攻击与3D稀疏和连续帧渲染,并在强制时间连续性的同时明确抑制特定于人的语义特征。实验表明AdvSerial在多个检测器上有高成功率和强可转移性,揭示了高角度监控下持续的、时间上一致的故障模式,为安全关键的基础设施部署设计运动感知和3D感知防御提供了动力。

英文摘要

AI-based visual perception systems are increasingly deployed in infrastructure surveillance, including roadside monitoring units, highway cameras, and smart-city pedestrian management systems. The security vulnerability of these systems to physical adversarial attacks poses a direct threat to the reliable operation of transportation infrastructure. We propose AdvSerial, a dynamic 2D--3D joint optimization framework for generating continuous high-angle physical adversarial patches against pedestrian detectors in infrastructure-based scenarios. We UV-map a boundary-aware quilted texture onto 3D garments, combine 2D digital attacks with 3D sparse- and continuous-frame rendering, and explicitly suppress person-specific semantic features while enforcing temporal continuity. A Feature Smooth Quilting strategy reduces visible patch boundaries and bounds cross-seam feature discontinuities. A serial-frame loss encourages long uninterrupted sequences of detection failures. In physical world experiments, AdvSerial achieves a 74.8% attack success rate on YOLO-v5 and degrades mean detection confidence from 84.30% to 39.38%. Experiments spanning eight detectors with different architectures demonstrate strong transferability. Notably, it achieves an $89.71%$ attack success rate on YOLO-v2 and resists both patch-detection defenses (NapGuard) and 3D-temporal perception (Sparse4D-v3). The results reveal persistent, temporally consistent failure modes under high-angle surveillance, and motivate the design of motion-aware and 3D-aware defenses for security-critical infrastructure deployments.

URL PDF HTML 收藏
2607.16828 2026-07-21 cs.CV 新提交

UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation

UniNDM:文本到图像生成中针对性内容的统一噪声驱动检测与缓解框架

Yao Huang, Yitong Sun, Huanran Chen, Ruochen Zhang, Shouwei Ruan, Ranjie Duan, Maoxun Yuan, Yinpeng Dong, Hui Xue, Xiaochun Cao, Xingxing Wei

机构 * Institute of Artificial Intelligence, State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室人工智能研究院) College of Artificial Intelligence, Tsinghua University(清华大学人工智能学院) Security Department, Alibaba Group(阿里巴巴集团安全部) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-Sen University(中山大学深圳校区网络科学与技术学院)

AI总结 针对文本到图像生成易受隐式性提示影响的问题,提出UniNDM统一噪声驱动框架。利用早期预测噪声的可分离性开发轻量级检测器,引入噪声增强自适应负引导缓解问题,扩展到扩散变压器架构,实验显示比现有方法有显著改进。

Comments 18 pages, 10 figures, accepted by TPAMI

详情
AI中文摘要

尽管文本到图像扩散模型具有强大的生成能力,但它们容易受到隐式性提示的影响,由于模型偏差或训练数据中的潜在相关性,微妙线索会意外生成不当内容。现有安全机制存在根本局限性。为此,我们提出UniNDM,一个统一的噪声驱动框架,通过扩散过程中的噪声动态来重新思考安全机制。我们发现早期预测噪声在正常和性明确内容之间具有内在可分离性,并理论证明其语义浓度随时间步长二次增加。利用此特性,我们开发了轻量级基于噪声的检测器,准确率高且几乎无计算开销。对于缓解,我们引入噪声增强自适应负引导,通过大语言模型动态生成特定上下文负提示,同时通过抑制对明确令牌的注意力集中来优化初始噪声。我们还将框架扩展到新兴的扩散变压器架构。综合实验表明,我们的方法比现有方法有显著改进。

英文摘要

Despite the impressive generative capabilities of text-to-image diffusion models, they remain vulnerable to implicit sexual prompts, where subtle cues disguised as benign terms or adversarial tokens unexpectedly generate the inappropriate content due to model biases or latent correlations in training data. Existing safety mechanisms face fundamental limitations: detection methods primarily identify explicit content and fail to capture implicit malicious intent, while mitigation approaches rely on static negative prompts inadequate for diverse implicit scenarios. To address these challenges, we propose UniNDM, a unified noise-driven framework that rethinks safety mechanisms through the lens of noise dynamics in diffusion processes. Our key insight is that early-stage predicted noise exhibits inherent separability between normal and sexually explicit content, which we theoretically demonstrates quadratically increasing semantic concentration with timestep. Leveraging this property, we develop a lightweight noise-based detector achieving superior accuracy with virtually no computational overhead. For mitigation, we introduce noise-enhanced adaptive negative guidance: dynamically generating context-specific negative prompts via large language models to handle diverse implicit content, while optimizing initial noise by suppressing attention concentration on explicit tokens to provide comprehensive protection. Besides the U-Net-based diffusion models, we further extend our framework to emerging Diffusion Transformer architectures through region-constrained semantic guidance tailored for their unified multimodal attention. Comprehensive experiments across U-Net models and DiT models on both natural and adversarial datasets demonstrate substantial improvements over state-of-the-art methods, including SLD, UCE, Safree, etc. Our code is publicly available at https://github.com/Aries-iai/UniNDM.

URL PDF HTML 收藏
2605.18920 2026-07-21 cs.IR cs.AI 版本更新

SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation

SynGR:释放跨模态协同在生成推荐中的潜力

Wei Chen, Xingyu Guo, Shuang Li, Fuwei Zhang, Meng Yuan, Jing Fan, Zhao Zhang, Deqing Wang, Fuzhen Zhuang

机构 * School of Artificial Intelligence, Beihang University, Beijing, China(北京航空航天大学人工智能学院) School of Computer Science and Engineering, Beihang University, Beijing, China(北京航空航天大学计算机科学与工程学院)

AI总结 本文提出SynGR框架,通过显式鼓励生成过程中的跨模态依赖,以捕捉新兴物品语义,从而提升生成推荐性能。

Comments Accepted by ICML 2026, 15 pages

详情
AI中文摘要

生成推荐(GR)通过将物品推荐问题建模为序列到序列生成任务,已成为一种有前景的范式。最近的研究将多模态信号纳入其中,以提供更丰富的token级证据。然而,现有方法主要依赖对齐中心融合,并未充分探索跨模态的协同信息。实际上,协同信息在捕捉无法从单一模态推断出的新兴物品属性中起着关键作用。这些属性编码了内在的物品语义并指导用户偏好,使模型能够超越表层特征匹配。为了解决这一限制,我们提出了SynGR,一种协同生成推荐框架,该框架在生成过程中显式鼓励利用跨模态依赖。通过限制对主导模态的过度依赖,SynGR使模型能够捕捉超出共享或模态特定信号的新兴物品语义。在三个基准数据集上的广泛实验表明,SynGR实现了优越的性能。

英文摘要

Generative Recommendation (GR) has emerged as a promising paradigm by formulating item recommendation as a sequence-to-sequence generation task over item identifiers. Recent studies have incorporated multimodal signals to provide richer token-level evidence for generation. However, existing approaches largely rely on alignment-centric fusion and underexplore synergistic information across modalities. In practice, synergistic information plays a critical role in capturing emergent item properties that cannot be inferred from any single modality alone. Such properties encode intrinsic item semantics and guide user preferences, enabling models to move beyond surface-level feature matching. To address this limitation, we propose \textbf{SynGR}, a synergistic generative recommendation framework that explicitly encourages the exploitation of cross-modal dependencies during generation. By constraining overreliance on dominant modalities, SynGR enables the model to capture emergent item semantics beyond shared or modality-specific signals. Extensive experiments across three benchmark datasets demonstrate that SynGR achieves superior performance.

URL PDF HTML 收藏
2604.25693 2026-07-21 cs.AI 版本更新

RADD: Retrieval-Augmented Discrete Diffusion for Multi-Modal Knowledge Graph Completion

RADD:基于检索的离散扩散用于多模态知识图谱补全

Guanglin Niu, Bo Li

机构 * School of Artificial Intelligence, Beihang University(北航人工智能学院)

AI总结 本文提出RADD框架,通过分离检索与重排序过程,提升多模态知识图谱补全的性能,实验表明其在多个基准上表现最佳。

Comments 13 pages, 3 figures, 8 tables

详情
AI中文摘要

大多数多模态知识图谱补全(MMKGC)模型使用一个嵌入评分器同时进行全实体集检索和最终决策。我们认为这种耦合是核心瓶颈:全局高召回率搜索和局部细粒度歧义消除需要不同的归纳偏置。因此,我们提出一种检索增强的离散扩散(RADD)框架,以解耦MMKGC中的检索与重排序。一个关系感知的多模态知识图谱嵌入(KGE)检索器同时充当全局检索器和蒸馏教师,而一个条件性离散去噪器则在短名单层面进行实体身份生成以进行重排序。训练结合了KGE监督、去噪交叉熵以及从检索器到去噪器的温度缩放蒸馏。在推理阶段,设计的Diff-Rerank首先通过检索器生成前K个短名单,然后通过去噪器对其进行重排序,确保召回是精度的严格前提。在三个MMKGC基准上的实验表明,RADD在性能和一致性方面均优于强大的单模态、多模态和LLM基线模型,而消融实验进一步验证了每个组件的贡献。

英文摘要

Most multi-modal knowledge graph completion (MMKGC) models use one embedding scorer to conduct both retrieval over the full entity set and final link prediction. We argue that this coupling is a core bottleneck: global high-recall search and local fine-grained disambiguation require different inductive biases. Therefore, we propose a Retrieval-Augmented Discrete Diffusion (RADD) framework to decouple retrieval and reranking for MMKGC. A relation-aware multimodal knowledge graph embedding (KGE) retriever serves as both global retriever and distillation teacher, while a conditional discrete denoiser performs shortlist-level entity-identity generation for reranking. Training combines KGE supervision, denoising cross-entropy, and temperature-scaled distillation from the retriever to the denoiser. At inference, the designed Diff-Rerank first forms a top-K shortlist with the retriever and then reranks it with the denoiser, ensuring that recall is a strict requirement for precision. Experiments on three MMKGC benchmarks show that RADD achieves the best performance and consistent gains over strong unimodal, multimodal, and LLM-based baselines, while ablations further verify each component's contribution.

URL PDF HTML 收藏
2604.02808 2026-07-21 cs.CV 版本更新

CMCC-ReID: Cross-Modality Clothing-Change Person Re-Identification

CMCC-ReID: 跨模态服装变化人员重识别

Haoxuan Xu, Hanzi Wang, Guanglin Niu

机构 * School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) School of Informatics, Xiamen University(厦门大学信息学院)

AI总结 本文提出CMCC-ReID任务,解决跨模态和服装变化的人员匹配问题,构建SYSU-CMCC基准数据集,并提出PIA网络,通过DBDL和BPL模块提升重识别性能。

Comments ECCV2026

详情
AI中文摘要

人员重识别(ReID)在长期监控场景中面临模态差异和服装变化的严峻挑战。现有研究在可见-红外ReID(VI-ReID)或服装变化ReID(CC-ReID)方面取得显著进展,但现实监控系统常同时面临这两种挑战。为解决这一被忽视但现实的问题,本文定义了新的任务——跨模态服装变化重识别(CMCC-ReID),旨在实现模态和服装变化下的人员匹配。为推进该方向的研究,本文构建了新的基准SYSU-CMCC,其中每个身份在可见和红外域中以不同着装出现,反映了长期监控的双重异质性。为解决CMCC-ReID问题,本文提出渐进身份对齐网络(PIA),逐步缓解服装变化和模态差异问题。具体而言,双分支解耦学习(DBDL)模块分离身份相关线索与服装相关因素,以实现服装无关的表示,而双向原型学习(BPL)模块在嵌入空间中进行内模态和外模态对比,以弥合模态差距并进一步抑制服装干扰。在SYSU-CMCC数据集上的大量实验表明,PIA为该新任务建立了强大的基准,并显著优于现有方法。

英文摘要

Person Re-Identification (ReID) faces severe challenges from modality discrepancy and clothing variation in long-term surveillance scenario. While existing studies have made significant progress in either Visible-Infrared ReID (VI-ReID) or Clothing-Change ReID (CC-ReID), real-world surveillance system often face both challenges simultaneously. To address this overlooked yet realistic problem, we define a new task, termed Cross-Modality Clothing-Change Re-Identification (CMCC-ReID), which targets pedestrian matching across variations in both modality and clothing. To advance research in this direction, we construct a new benchmark SYSU-CMCC, where each identity is captured in both visible and infrared domains with distinct outfits, reflecting the dual heterogeneity of long-term surveillance. To tackle CMCC-ReID, we propose a Progressive Identity Alignment Network (PIA) that progressively mitigates the issues of clothing variation and modality discrepancy. Specifically, a Dual-Branch Disentangling Learning (DBDL) module separates identity-related cues from clothing-related factors to achieve clothing-agnostic representation, and a Bi-Directional Prototype Learning (BPL) module performs intra-modality and inter-modality contrast in the embedding space to bridge the modality gap while further suppressing clothing interference. Extensive experiments on the SYSU-CMCC dataset demonstrate that PIA establishes a strong baseline for this new task and significantly outperforms existing methods.

URL PDF HTML 收藏
2607.16074 2026-07-20 cs.DC cs.AI cs.SE 新提交

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

JoyNexus:面向服务的视觉语言动作模型多租户训练后处理

Haoran Sun, Wentao Zhang, Junyang Hua, Hedan Yang, Yongjian Guo, Yifei Zhang, Xiaolong Xiang, Mingxi Luo, Jing Long, Chen Zhao, Chen Zhou, Wanting Xu, Qiming Yang, Hui Zhang, Song Wang, Xiaodong Bai, Shuai Di, Xu Chu, Xiaotie Deng, Yicheng Gong, Junwu Xiong

机构 * Peking University(北京大学) Beihang University(北航) Beijing Institute of Technology(北京理工大学) Tsinghua University(清华大学)

AI总结 针对VLA模型训练后处理问题,提出JoyNexus统一服务,解耦相关服务,多租户可通过API调用,引入组批处理提高效率,经评估能减少GPU时间,提升服务利用率。

Comments 23 pages, 12 figures

详情
AI中文摘要

视觉语言动作(VLA)模型的训练后处理至关重要。现有计算服务通常为单个租户分配专用的GPU和CPU资源,存在基础设施适配负担重、计费模式不合理等问题。为此提出JoyNexus,它解耦了训练模型服务、推理模型服务和环境服务,多租户可通过API调用。还引入组批处理提高效率,通过工作负载模拟和组批处理管道评估,结果表明其能减少GPU时间并提高服务利用率。

英文摘要

The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloads both expensive for tenants and inefficient for the service provider. To address these challenges, we present JoyNexus, a unified service for multi-tenant VLA supervised fine-tuning, reinforcement learning, and evaluation. JoyNexus decouples the Training Model Service, Inference Model Service, and Environment Service, each accessed through APIs and backed by resident shared base models with tenant-specific slots. Tenants can directly invoke high-level semantic APIs for training, rollout, and evaluation, or compose custom algorithms using lower-level APIs and their assigned endpoints. Multiple tenants submit workloads concurrently; their action modules, optimizers, rollout records, and policy versions remain isolated, and the service is scheduled by the global Training Queue and Inference Queue. To further improve multi-tenant training efficiency, JoyNexus introduces group batching for heterogeneous VLA data schemas that share a compatible model-facing prefix, enabling a single shared backbone forward pass over grouped samples. Finally, we evaluate JoyNexus through workload simulation and a group-batching pipeline in a realistic embodied scenario. Results show that, compared with isolated single-tenant execution, JoyNexus reduces aggregate GPU time and improves service utilization via cross-tenant scheduling on shared resources.

URL PDF HTML 收藏
2607.15079 2026-07-20 cs.AI 版本更新

BrainPilot: Automating Brain Discovery with Agentic Research

BrainPilot:通过智能研究实现大脑发现自动化

Haoxuan Li, Tianci Gao, Jianhe Li, Yang Fan, Runze Shi, Weiran Wang, Tianxiang Zhao, Zezhao Wu, Xiaoyang Jiang, Qihui Zhang, Jia Li, Xiao Xiao, Kai Du, Xiaoxuan Jia, Chao Xie, Lu Mi

机构 * College of AI, Tsinghua University(清华大学人工智能学院) Shanghai Qizhi Institute(上海期智研究院) Business School, Renmin University of China(中国人民大学商学院) School of Physics, Beihang University(北京航空航天大学物理学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) Behavioral and Cognitive Neuroscience Center, Institute of Science and Technology for Brain-Inspired Intelligence, Fudan University(复旦大学脑科学与智能技术研究院行为与认知神经科学中心) College of Engineering, Georgia Institute of Technology(佐治亚理工学院工程学院) School of Life Sciences & IDG/McGovern Institute for Brain Research, Tsinghua University(清华大学生命科学学院&清华-IDG/麦戈文脑科学研究院) School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院) Weixian College, Tsinghua University(清华大学未央学院) Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理学与认知科学系)

AI总结 研究针对脑科学研究整合证据难、人工智能代理有缺陷的问题,提出完全开源的多智能体系统BrainPilot,它有可追溯日志和验证结果,含知识库与技能库,经实验评估,其开源模型以低成本达先进框架性能。

详情
AI中文摘要

理解大脑越来越依赖跨尺度、模态和学科整合证据。解决单个研究问题需要一系列协调操作。人工智能代理有望加速这一过程,但当前代理在脑科学领域缺乏专业知识,可能编造主张,在多步推理中偏离,且专家干预点少。我们提出了BrainPilot,一个完全开源的多智能体系统,它通过可追溯的日志和经智能体验证的结果加速脑科学研究。主要研究者(PI)智能体协调基于精心策划的领域知识的专家智能体,包括一个包含7233个索引条目的统一脑科学知识库和一个涵盖七个研究领域的72个可重复使用方法单元的技能库。每个主要步骤都记录在追踪图中,审核智能体将伪造检查集成到工作流程中。为了评估,我们运行了来自智能体期末考试的三个脑科学任务,引入了自己的基准BrainPilotBench-v0,并展示了其他端到端案例研究。在这些评估中,具有开源主干模型的BrainPilot以更低成本实现了与最先进智能体框架相当的性能。

英文摘要

Understanding the brain increasingly depends on integrating evidence across scales, modalities, and disciplines. Addressing a single research question therefore requires a coordinated sequence of operations, from surveying prior work to executing analyses and interpreting results in light of domain knowledge. AI agents promise to accelerate this process, but current agents lack domain expertise in brain science, may fabricate claims, drift during multi-step reasoning, and offer few defined points for expert intervention. These failures are especially costly in brain science, where conclusions feed into downstream scientific claims and depend on laboratory-specific expertise and careful human judgment. We present \textbf{BrainPilot} a \textbf{fully open-source} multi-agent system that accelerates brain science research with traceable logs and agent-verified results. A principal investigator (PI) agent coordinates specialist agents grounded in curated domain knowledge: a unified brain science knowledge base containing 7{,}233 indexed items and a skill library of 72 reusable methodology units across seven research domains. Every major step is recorded in the Graph of Trace, an auditable record that links subgoals, tool use, evidence, and claims and allows researchers to follow and inspect the workflow. An Auditor agent further integrates fabrication checking into the workflow. For evaluation, we run three brain science tasks from Agents' Last Exam, introduce our own benchmark, \textbf{BrainPilotBench-v0}, and present additional end-to-end case studies. Across these evaluations, BrainPilot with an open-source backbone model attains performance comparable to state-of-the-art agent framework with less costs.

URL PDF HTML 收藏
2502.06818 2026-07-20 cs.LG 版本更新

Rethinking the Global Knowledge of CLIP in Training-Free Open-Vocabulary Semantic Segmentation

重新思考CLIP的全局知识以实现无训练开放词汇语义分割

Jingyun Wang, Cilin Yan, Guoliang Kang

机构 * Beihang University(北航大学)

AI总结 本文提出GCLIP,通过重塑最后块注意力和Value嵌入来提取CLIP的全局知识,提升无训练开放词汇语义分割性能。

Comments TMM 2026

详情
AI中文摘要

近期工作通过修改CLIP实现无训练开放词汇语义分割(TF-OVSS)。在vanilla CLIP中,基于块的图像表示主要编码同质的图像级属性,阻碍了CLIP在密集预测任务中的应用。先前TF-OVSS工作牺牲全局性以增强CLIP特征的局部性,通过使每个块主要关注自身或其邻近块内的狭窄局部窗口。经过修改后,CLIP聚合全局上下文信息的能力大幅减弱。本文重新思考CLIP编码的全局知识,提出GCLIP,探讨如何提取和利用CLIP的有益全局知识以实现TF-OVSS。由于每个块的表示最终由注意力权重和Value嵌入决定,我们提出重塑最后块的注意力和Value嵌入以将有用的全局上下文聚合到最终特征中。首先,我们旨在为最后块的注意力赋予图像级属性,而不引入跨块的同质注意力模式。为此,我们融合来自全局令牌出现块的注意力与Query-Query注意力。其次,我们旨在使最后块注意力模块的Value嵌入更具语义相关性。为此,我们设计了一种新的通道抑制策略。在五个标准基准上的广泛实验表明,我们的方法在性能上始终优于先前的最先进方法。

英文摘要

Recent works modify CLIP to perform open-vocabulary semantic segmentation in a training-free manner (TF-OVSS). In vanilla CLIP, patch-wise image representations mainly encode homogeneous image-level properties, which hinders the application of CLIP to the dense prediction task. Previous TF-OVSS works sacrifice globality to enhance the locality of CLIP features, by making each patch mainly attend to itself or its neighboring patches within a narrow local window. With their modifications,the ability of CLIP to aggregate global context information is largely weakened. Differently, in this paper, we rethink the global knowledge encoded by CLIP and propose GCLIP to answer how to extract and utilize beneficial global knowledge of CLIP for TF-OVSS. As the representation of each patch is finally determined by the attention weights and the Value embeddings, we propose to reshape the last-block attention and Value embeddings to aggregate useful global context into final features. Firstly, we aim to equip the last-block attention with image-level properties while not introducing homogeneous attention patterns across patches. To realize the goal, we fuse the attention from the global-token emerging blocks with the Query-Query attention. Secondly, we aim to make Value embeddings of the last-block attention module more semantically correlated. To realize this, we design a novel channel suppression strategy.Extensive experiments on five standard benchmarks demonstrate that our method consistently outperforms previous state-of-the-arts.

URL PDF HTML 收藏
2602.01244 2026-07-20 cs.CL 版本更新

Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments

从Docker化环境生成大规模终端代理轨迹

Siwei Wu, Yizhi Li, Yuyang Song, Wei Zhang, Yang Wang, Riza Batista-Navarro, Xian Yang, Mingjie Tang, Bryan Dai, Jian Yang, Chenghua Lin

机构 * University of Manchester(曼彻斯特大学) IQuest Research(IQuest研究) Beihang University(北航大学) Sichuan University(四川大学) Multimodal Art Projection Research Community(多模态艺术投影研究社区)

AI总结 本文提出TerminalTraj,通过Docker化环境生成高质量终端轨迹,提升代理模型在终端任务中的性能表现。

Comments Accepted as a Spotlight paper at ICML 2026

详情
AI中文摘要

训练基于终端的任务的代理模型严重依赖于高质量的终端轨迹,这些轨迹能够捕捉跨不同领域的长周期交互。然而,由于两个关键要求,大规模构建此类数据仍然具有挑战性:可执行性,因为每个实例都需要合适的且通常不同的Docker环境;以及可验证性,因为异构任务输出排除了统一的标准验证。为了解决这些挑战,我们提出了TerminalTraj,一个可扩展的流水线,它(i)过滤高质量的仓库以构建Docker化的执行环境,(ii)生成与Docker对齐的任务实例,(iii)合成具有可执行验证代码的代理轨迹。使用TerminalTraj,我们整理了32,000个Docker镜像,并在八个领域生成了50,733个经过验证的终端轨迹。在该数据上训练的模型,使用Qwen2.5-Coder作为骨干,能够在TerminalBench(TB)上实现一致的性能提升,在TB~1.0上提升高达20%,在TB~2.0上提升10%。值得注意的是,TerminalTraj-32B在参数少于100B的模型中表现强劲,达到TB~1.0的35.30%和TB~2.0的22.00%,并在测试时表现出改进的扩展行为。所有代码和数据均在https://github.com/Wusiwei0410/TerminalTraj上提供。

英文摘要

Training agentic models for terminal-based tasks critically depends on high-quality terminal trajectories that capture realistic long-horizon interactions across diverse domains. However, constructing such data at scale remains challenging due to two key requirements: \textbf{\emph{Executability}}, since each instance requires a suitable and often distinct Docker environment; and \textbf{\emph{Verifiability}}, because heterogeneous task outputs preclude unified, standardized verification. To address these challenges, we propose \textbf{TerminalTraj}, a scalable pipeline that (i) filters high-quality repositories to construct Dockerized execution environments, (ii) generates Docker-aligned task instances, and (iii) synthesizes agent trajectories with executable validation code. Using TerminalTraj, we curate 32K Docker images and generate 50,733 verified terminal trajectories across eight domains. Models trained on this data with the Qwen2.5-Coder backbone achieve consistent performance improvements on TerminalBench (TB), with gains of up to 20\% on TB~1.0 and 10\% on TB~2.0 over their respective backbones. Notably, \textbf{TerminalTraj-32B} achieves strong performance among models with fewer than 100B parameters, reaching 35.30\% on TB~1.0 and 22.00\% on TB~2.0, and demonstrates improved test-time scaling behavior. All code and data are available at https://github.com/Wusiwei0410/TerminalTraj.

URL PDF HTML 收藏
2607.15278 2026-07-17 cs.CV 新提交

Hierarchical Denoising For Multi-Step Visual Reasoning

用于多步视觉推理的分层去噪

Zezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou, Ruibin Yuan, Yuhan Rui, Hengzhe Sun, Zhuoqun Wu, Yuming Li, Siyuan Qian, Sirui Han, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机科学学院多媒体信息处理技术国家重点实验室) The Hong Kong University of Science and Technology(香港科技大学) Beihang University(北京航空航天大学) Fuzhou University(福州大学) Muka Robotics(木卡机器人)

AI总结 研究针对视频模型多步推理不足的问题,提出HDR框架,通过树形层次结构和稀疏分层注意力模式进行多步推理,在新基准测试中提升了成功率和进度,推理更快,数据效率更高,在机器人实验中展现潜力。

详情
AI中文摘要

视频模型正在演变为视觉基础模型,但仍缺乏类似人类的多步推理能力。流式自回归扩散模型高效但推理有限,双向扩散虽能全局修正但推理成本高。我们提出了HDR,一个将分层潜在因素集成到因果视频生成中进行多步推理的统一框架。HDR将视频潜在因素组织成树形层次结构,在流式输出前实现从粗到细的推理。粗去噪层保留不确定假设用于全局规划,细层逐步将其细化为具体视觉状态。稀疏分层注意力模式进一步降低时间注意力成本。我们引入了一个具有分布外情况的分层多步视频推理基准,涵盖六个任务。与流式自回归扩散基线相比,HDR将成功率从34.22提高到60.29,平均进度从76.00提高到89.56,推理速度比双向扩散快54.2倍,在仅2%训练数据时保留82.9%的全数据性能。真实世界机器人实验进一步证明了HDR在物理交互和世界建模方面的潜力。

英文摘要

Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradigms struggle to achieve logical consistency and low-latency streaming for complex reasoning tasks. We propose HDR (Hierarchical Denoising for Visual Reasoning), a unified framework that integrates hierarchical latents into causal video generation for multi-step reasoning. HDR organizes video latents into a tree-structured hierarchy, enabling coarse-to-fine reasoning before streaming output. Coarse denoising layers preserve uncertain hypotheses for global planning, while finer layers progressively refine them into concrete visual states. A sparse hierarchical attention pattern (SHAP) further reduces temporal attention costs. We introduce a level-stratified multi-step video reasoning benchmark with out-of-distribution cases, covering six tasks: maze navigation, Tower of Hanoi, one-line drawing, sliding puzzle, Sokoban, and water pouring. Compared with streaming autoregressive diffusion baselines, HDR improves success from 34.22 to 60.29 (76.2% relative gain) and increases average progress from 76.00 to 89.56, demonstrating more consistent reasoning trajectories. HDR maintains low-latency streaming at 0.70 seconds per latent, achieving 54.2 times faster inference than bidirectional diffusion. It also retains 82.9% of full-data performance with only 2% training data, compared with 52.0% for bidirectional diffusion. Real-world robot experiments further demonstrate HDR's potential for physical interaction and world modeling. Project demo: https://hierarchical-diffusion-reasoning.github.io/.

URL PDF HTML 收藏
2607.14510 2026-07-17 cs.AI 新提交

VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence

VLT:用于工业智能的视觉-语言-时间序列多模态基础模型

Haiteng Wang, Jingheng Yan, Xiaokang Wang, Lei Ren

机构 * School of Automation Science and Electrical Engineering, Beihang University(北京航空航天大学自动化科学与电气工程学院) Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) State Key Laboratory of Intelligent Manufacturing System Technology(智能制造系统技术国家重点实验室) School of Computer Science and Artificial Intelligence, Zhengzhou University(郑州大学计算机科学与人工智能学院)

AI总结 针对工业时间序列单模态建模局限及连接时间序列与文本语义的挑战,提出VLT多模态基础模型,通过设计Time-MoE等机制联合建模多种模态,经实验验证其在多种复杂设置下优于现有方法,提升了鲁棒性和泛化能力。

Comments 18 pages, 13 figures, and 13 tables, including supplementary material. Haiteng Wang and Jingheng Yan contributed equally to this work

详情
AI中文摘要

工业时间序列是预测与健康管理(PHM)的基础,可确保航空发动机等工业设备的可靠性和安全性。但现有方法多限于单模态建模,限制了其在复杂场景中的泛化。尽管大语言模型的进展为多模态学习带来新机遇,但连接连续时间序列信号和离散文本语义仍是挑战。为此提出VLT,一个联合建模时间序列、频谱视觉表示和文本知识的多模态基础模型。关键在于利用频谱作为视觉桥梁连接连续时间信号和离散语义。具体设计了时间感知专家混合模型(Time-MoE)捕捉异构时间动态,频率-文本增强学习者在共享表示空间中联合建模频谱和语义特征。还引入以时间为中心的梯度对齐机制减轻跨模态优化冲突。在多个工业数据集上的大量实验表明,VLT优于现有方法,在少样本、有噪声和不完全模态设置下具有卓越的鲁棒性和泛化能力。

英文摘要

Industrial time series serve as the foundation for Prognostics and Health Management (PHM) to ensure the reliability and safety of industrial equipment such as aero-engines. However, existing approaches are typically limited to single-modality modeling, which restricts their generalization in complex scenarios. Although recent advances in large language models (LLMs) provide new opportunities for multimodal learning, bridging continuous time-series signals and discrete textual semantics remains an open challenge. To this end, we propose VLT, a multimodal foundation model that jointly models time-series, frequency-spectrum visual representations, and textual knowledge. A key insight is to utilize the frequency spectrum as a visual bridge to connect continuous temporal signals with discrete semantics. Specifically, a Time-aware Mixture-of-Experts (Time-MoE) is designed to capture heterogeneous temporal dynamics, while a Frequency-Text Augmented Learner enables joint modeling of spectral and semantic features within a shared representation space. Furthermore, a time-centric gradient alignment mechanism is introduced to mitigate cross-modal optimization conflicts via gradient normalization and reliability-aware dynamic reweighting. Extensive experiments on multiple industrial datasets demonstrate that VLT outperforms state-of-the-art methods, achieving superior robustness and generalization under few-shot, noisy, and incomplete-modality settings.

URL PDF HTML 收藏
2505.10999 2026-07-17 cs.CV 版本更新

Conditioning Residuals for Diffusion Models via Representation Feedback

通过表示反馈对扩散模型的残差进行条件设定

Weilai Xiang, Hongyu Yang, Di Huang, Yunhong Wang

机构 * State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(虚拟现实技术与系统国家重点实验室,北京航空航天大学) Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)

AI总结 研究扩散模型中能否利用固有路径路由内部推断语义,提出条件残差这一轻量级反馈机制,通过反馈紧凑特征摘要提供自适应指导,实验显示其提升了生成性能及下游任务表现,还揭示了训练动态和特征结构的改进。

Comments Substantially revised and retitled. Code available at https://github.com/FutureXiang/ddae_plus_plus

详情
AI中文摘要

扩散模型如今是多媒体生成的常见基础,其生成训练中会出现有用的中间表示。然而标准架构通过主特征流传播这些表示,未将编码语义明确重新引入后续去噪层。同时,此类主干已为预定义输入的全局调制提供了条件设定路径。本文研究该固有路径能否将内部推断语义作为不断演变的、依赖样本的线索进行路由。我们提出条件残差,一种轻量级反馈机制,将聚合特征转换为添加到条件嵌入的残差。通过反馈紧凑特征摘要,它提供自适应生成指导并促进更紧的语义瓶颈,无需外部编码器、辅助目标或采样时间更改。实验表明在生成性能上有持续提升,且下游线性探测和分割中有更强表示。机理分析揭示了改进的生成训练动态和重塑的特征结构。

英文摘要

Diffusion models now serve as a common foundation for multimedia generation, and useful intermediate representations emerge during their generative training. Standard architectures, however, propagate these representations through the main feature stream, without explicitly reintroducing their encoded semantics to later denoising layers. Meanwhile, such backbones already provide a conditioning pathway for global modulation by predefined inputs. This work examines whether this native pathway can also route internally inferred semantics as evolving, sample-dependent cues. We propose Conditioning Residuals, a lightweight feedback mechanism that converts aggregated features into residuals added to condition embeddings. By feeding back compact feature summaries, it provides adaptive generative guidance and encourages a tighter semantic bottleneck, without external encoders, auxiliary objectives, or sampling-time changes. It supports feedback at one or multiple depths in UNet and DiT backbones, with negligible overhead. Across diffusion formulations, backbone configurations, and datasets, experiments show consistent gains in generative performance, along with stronger representations in downstream linear probing and segmentation. Mechanistic analyses reveal improved generative training dynamics and reshaped feature structure, suggesting a grounded, generalizable way to enhance diffusion backbones from within.

URL PDF HTML 收藏
2607.13101 2026-07-16 cs.LG cs.AI 新提交

TSSM: Triaxial State Space Model for Global Station Weather Forecasting with Temporal-Variable-Historical Modeling

TSSM:用于全球台站天气预报的具有时间可变历史建模的三轴状态空间模型

Songru Yang, Zili Liu, Tao Han, Ben Fei, Fenghua Ling, Lei Bai, Chang Liu, Xiangyang Ji, Zhenwei Shi, Zhengxia Zou

机构 * Beihang University(北京航空航天大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学)

AI总结 研究针对全球台站天气预报中现有方法的局限,提出三轴状态空间模型TSSM,通过时间可变历史范式纳入历史数据,设计扫描捕捉多种特性,在多数据集上性能优异,尤其在长期和迭代预测及应对缺失观测方面优势明显。

详情
AI中文摘要

全球台站天气预报对于关键区域的局部和极端天气预报至关重要。尽管努力利用回溯窗口,但现有方法在准确性提升方面有限,且在极端事件和误差累积方面存在困难。这些限制源于对短期模式的过度依赖,不足以捕捉混沌天气动态。为解决此问题,我们提出了一种新颖的三轴状态空间模型(TSSM),具有历史增强的时间可变历史范式,纳入周期对齐的历史天气数据以补偿时间回溯窗口之外的长期、大规模周期性和全窗口天气模式。具体而言,TSSM将历史样本堆叠成周期对齐的批次,预测由历史和当前观测因果支持。设计了时间、可变和历史扫描以捕捉轴向时间依赖性、可变相关性和历史演变。这种结构分层共享以对季节性到极端事件进行建模,同时减轻历史模式之间的错位。TSSM在Weather-5K(迄今为止最大的台站天气数据集)上实现了SOTA性能,在准确性和极端事件指标上分别提高了10%和61%,在人工参与的数据集上获得了95%的最佳或次佳结果。其优势在长期和迭代预测中更为明显,在240小时时增益达到37.5%,在48小时×5次迭代设置下高达103.5%。此外,与基线的<43%相比,TSSM在高达80%的缺失观测下仍保持>90%的性能,展示了在全球原位观测网络中进行可靠全球台站天气预报的稳健性和实际潜力。

英文摘要

Global Station Weather Forecasting (GSWF) is pivotal for localized and extreme weather prediction over key regions. Despite efforts to exploit look-back windows, existing methods show limited accuracy gains and struggle with extreme events and error accumulation. These limitations stem from overreliance on short-term patterns, which are insufficient to capture chaotic weather dynamics, especially under partial observations. To address this problem, we propose a novel Triaxial State Space Model (TSSM) with a history-enhanced Temporal-VariableHistorical paradigm, which incorporates period-aligned historical weather data to compensate for long-term, large-scale periodic, and full-window weather patterns beyond the temporal lookback window. Specifically, TSSM stacks historical samples into period-aligned batches, where forecasting is causally supported by historical and current observations. Temporal, variable, and historical scanning are designed to capture axial temporal dependencies, variable correlations, and historical evolution. This structure is hierarchically shared to model seasonal to extreme events while alleviating misalignment across historical patterns. TSSM achieves SOTA performance on Weather-5K, the largest station weather dataset to date, with 10% and 61% gains in accuracy and extreme event metrics, and obtains 95% best or second-best results on human-involved datasets. Its advantages are more pronounced in long-horizon and iterative forecasting, reaching a 37.5% gain at 240h and up to 103.5% under a 48h times 5 iterative setting. Moreover, TSSM retains > 90% performance under up to 80% missing observations, compared with < 43% for baselines, demonstrating robustness and practical potential for reliable GSWF in global in-situ observation networks.

URL PDF HTML 收藏
2607.12680 2026-07-15 cs.CV 新提交

ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning

ReflectVLN:通过反思推理训练视觉语言导航智能体

Jiahang Wang, Yirong Yang, Yanqing Zhu, Minghua Luo, Shichao Xie, Fei Liu, Mu Xu

机构 * Amap, Alibaba Group(高德软件有限公司,阿里巴巴集团) Beihang University(北京航空航天大学)

AI总结 研究针对现有视觉语言导航方法缺乏闭环机制问题,提出ReflectVLN框架,通过双向交互智能体决策,引入行动思维链训练方案,实验证明该框架在有限数据下提升成功率和路径效率,且具良好训练成本与可解释性。

详情
AI中文摘要

现有视觉语言导航方法常将视觉语言模型(VLM)与航点解码器结合生成多步行动计划,但缺乏明确闭环机制来跟踪语义进展、诊断执行失败及从长期导航中的错误积累中恢复。为填补这一空白,我们提出ReflectVLN,一个通过双向交互意图和执行智能体组织决策的智能体视觉语言导航框架。意图智能体执行子任务分解和反思,生成可执行的子任务描述作为纠正计划。执行智能体依据这些描述在当前观察下将其转化为短期行动,同时监测子目标进展并检测偏离行为。关键的是,ReflectVLN实现了闭环双向通信。为鼓励具有可解释中间推理的时间连贯决策,我们引入行动思维链(Action-CoT),一种用于行动生成的路径条件双查询训练方案。在标准视觉语言导航基准测试中的实验表明,ReflectVLN在有限数据预算下提高了成功率和路径效率,具有良好的训练成本,推理时高级意图调用更少,同时提供可解释的中间决策用于分析和协作。

英文摘要

Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution failures, and recovering from error accumulation in long-horizon navigation. To address this gap, we propose ReflectVLN, an agentic VLN framework that organizes decision-making through bidirectionally interactive intention and execution agents. The intention agent performs subtask decomposition and reflection, generating executable subtask descriptions as corrective plans. Conditioned on these descriptions, the execution agent grounds them into short-horizon actions under current observations while monitoring sub-goal progress and detecting off-track behavior. Crucially, ReflectVLN enables closed-loop bidirectional communication: the execution agent emits progress and deviation signals to trigger reflection and subtask updates on demand, and the intention agent returns structured guidance that reconditions subsequent actions for recovery. To encourage temporally coherent decisions with interpretable intermediate rationales, we introduce Action Chain-of-Thought (Action-CoT), a path-conditioned dual-query training scheme for action generation. Experiments on standard VLN benchmarks show that ReflectVLN improves success rates and path efficiency under a constrained data budget, with favorable training cost and fewer high-level intention calls at inference time, while providing interpretable intermediate decisions for analysis and collaboration. Code is available at: https://github.com/AIprogrammer/ReflectVLN

URL PDF HTML 收藏
2607.12678 2026-07-15 cs.CV cs.AI 新提交

Text-Aided Multi-Modal Panoptic Symbol Spotting for CAD Floor Plan Drawings

用于CAD平面图的文本辅助多模态全景符号识别

Yan Gong, Bohao Li, Bowen Du, Junchen Ye

机构 * CCSE Lab, Beihang University(北京航空航天大学CCSE实验室) School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院) The Hong Kong Polytechnic University(香港理工大学)

AI总结 针对CAD平面图全景符号识别中现有方法未充分利用文本注释的问题,提出多模态框架TextCAD,设计TACE编码注释语义,引入语义层次对齐框架,实验表明该方法有效提升符号识别性能并达最优。

详情
AI中文摘要

计算机辅助设计(CAD)平面图包含图形原语和文本注释,为智能设计理解提供互补的几何和语义线索。在CAD分析任务中,随着工业数字化和基于深度学习的自动化需求的增长,全景符号识别变得越来越重要。然而,大多数现有方法主要以原语为中心,未充分利用文本注释。即使是少数文本感知方法也常常只是表面处理注释,没有正确建模CAD注释的复杂语法和层次语义,导致语义丢失和次优的识别性能。为解决这些限制,我们提出TextCAD,一个联合建模图形原语和文本注释用于全景符号识别的多模态框架。具体而言,我们设计了一个类型-属性相关编码器(TACE),通过联合建模注释的类型和属性来明确编码注释中的组合语义。我们还引入了一个带有多级语义过滤(MSF)和原语下采样的语义层次对齐框架,它能在不同语义层次上自适应地将注释语义与图形原语对齐,并实现准确的跨模态语义注入和融合。在真实世界建筑设计数据集上的实验表明,TextCAD有效提高了符号识别性能并取得了当前最优的结果。

英文摘要

Computer-Aided Design (CAD) floor plan drawings contain both graphical primitives and textual annotations, which provide complementary geometric and semantic cues for intelligent design understanding. Among CAD analysis tasks, panoptic symbol spotting has become increasingly important with the growing demand for industrial digitalization and deep learning-based automation. However, most existing methods remain primarily primitive-centric and underexploit textual annotations, despite their critical semantic value. Even the few text-aware approaches often treat annotations only superficially, without properly modeling complex syntax and hierarchical semantics of CAD annotations, which leads to semantic loss and suboptimal spotting performance. To address these limitations, we propose TextCAD, a multimodal framework that jointly models graphical primitives and textual annotations for panoptic symbol spotting. Specifically, we design a Type-Attribute Correlation Encoder (TACE) to explicitly encode the compositional semantics within annotations by jointly modeling their types and attributes. We further introduce a Semantic Hierarchy Alignment framework with Multi-level Semantic Filtering (MSF) and primitive downsampling, which adaptively aligns annotation semantics with graphical primitives at different semantic levels and enables accurate cross-modal semantic injection and fusion. Experiments on real-world building-design datasets show that TextCAD effectively improves symbol spotting performance and achieves state-of-the-art results.

URL PDF HTML 收藏
2607.11245 2026-07-15 cs.SE cs.AI 版本更新

An Empirical Study for Android-to-OpenHarmony GUI Test Migration

从安卓到开源鸿蒙系统的图形用户界面测试迁移实证研究

Yakun Zhang, Xinjia Chen, Yiyun Chen, Yuxia Zhang, Mingyi Zhou, Xiang Gao, Shaokun Zhang, Li Li, Yunming Ye

机构 * Harbin Institute of Technology(哈尔滨工业大学) Beijing Institute of Technology(北京理工大学) Beihang University(北航) Peking University(北京大学)

AI总结 研究从安卓到开源鸿蒙系统的图形用户界面测试迁移问题,构建数据集,选择并适配两种先进迁移方法进行评估,发现现有方法效果不佳,进而提出增强方法ITeM-HM,显著提升了测试迁移成功率。

详情
AI中文摘要

为减少从安卓到开源鸿蒙系统测试相应应用所需的大量工程工作量,迁移现有图形用户界面测试用例成为关键问题。当前研究未针对开源鸿蒙提出解决方案,也未对该系统上的迁移方法进行系统评估。本文首次对从安卓到开源鸿蒙的测试迁移进行系统实证研究。构建了ATH基准数据集,选择两种先进测试迁移方法并适配到开源鸿蒙,从测试性能、失败根源和开源鸿蒙特性影响三个角度评估。结果显示现有方法在该场景效果不佳,通过分析失败案例提出基于ITeM的增强方法ITeM-HM,相对原ITeM成功率提升214%。

英文摘要

To reduce the substantial engineering effort required to test the corresponding applications from Android to OpenHarmony, migrating existing GUI test cases has become a critical problem. However, current research neither proposes solutions tailored for OpenHarmony nor provides a systematic evaluation of migration approaches on this system, leaving developers with limited empirical guidance in practice. In this paper, we present the first systematic empirical study of test migration from Android to OpenHarmony. Specifically, we first construct a dataset referred to as the ATH Benchmark, comprising 36 commercial applications with an average of over 9 billion downloads, along with 108 manually designed test cases. Second, we select two state-of-the-art test migration approaches (i.e., ReSPlay and ITeM) and adapt these two approaches to enable their execution on OpenHarmony. Third, we use the preceding infrastructure to evaluate these two approaches from three perspectives, including testing performance, root causes of failures, and the impact of OpenHarmony characteristics. Our results reveal that existing test migration approaches are less effective (15% success-rate on ReSPlay and 26% success-rate on ITeM) in Android-to-OpenHarmony scenarios. Through an in-depth analysis of failed cases, we identify that test performance is primarily hindered by OpenHarmony-specific characteristics, including technical architecture differences and unique ecosystem traits. Utilizing these findings, we propose an enhanced approach based on ITeM, referred as ITeM-HM, which incorporates specific OpenHarmony system features. As a result, ITeM-HM successfully achieves a 214% success-rate relative improvement over the original ITeM (from 26% to 81%).

URL PDF HTML 收藏
2605.09977 2026-07-15 cs.CV 版本更新

INFANiTE: Implicit Neural representation for high-resolution Fetal brain spatio-temporal Atlas learNing from clinical Thick-slicE MRI

INFANiTE:隐式神经表示用于高分辨率胎儿脑空间-时间大体图谱学习从临床厚切片MRI

Xiaotian Hu, Mingxuan Liu, Hongjia Yang, Tongxi Song, Yijin Li, Yifei Chen, Haoxiang Li, Zihan Li, Yingqi Hao, Ziyu Li, Yi Liao, Haibo Qu, Qiyuan Tian

机构 * Beihang University(北航大学) Tsinghua University(清华大学) Sichuan University(四川大学) University of Oxford(牛津大学)

AI总结 INFANiTE通过隐式神经表示方法,实现了从厚切片MRI中高效生成高分辨率胎儿脑空间-时间大体图谱,显著加快了图谱构建过程,提升了图谱的一致性、参考保真度和生物合理性。

详情
AI中文摘要

空间-时间胎儿脑图谱对于表征正常神经发育和识别先天异常至关重要。然而,现有图谱构建流程需要数天进行切片到体积重建(SVR)以生成高分辨率3D脑体积,并需要额外数天进行迭代体积配准,使从大规模群体中构建图谱变得不切实际。我们通过INFANiTE,一种隐式神经表示(INR)框架,实现了从临床厚切片MRI扫描中学习高分辨率胎儿脑空间-时间图谱,完全 bypass 了昂贵的SVR和迭代非刚性配准步骤,从而显著加速了图谱构建。广泛实验表明,INFANiTE在主体一致性、参考保真度、内在质量和生物合理性方面均优于现有基线方法,即使在具有挑战性的稀疏数据设置下也是如此。此外,INFANiTE将端到端处理时间(即从原始扫描到最终图谱)从数天减少到数小时,相比传统基于3D体积的流程(如SyGN),促进了大规模群体水平的胎儿脑分析。我们的代码在:https://anonymous.4open.science/r/INFANiTE-5D74 公开可用。

英文摘要

Spatio-temporal fetal brain atlases are important for characterizing normative neurodevelopment and identifying congenital anomalies. However, existing atlas construction pipelines necessitate days for slice-to-volume reconstruction (SVR) to generate high-resolution 3D brain volumes and several additional days for iterative volume registration, thereby rendering atlas construction from large-scale cohorts prohibitively impractical. We address these limitations with INFANiTE, an Implicit Neural Representation (INR) framework for high-resolution Fetal brain spatio-temporal Atlas learNing from clinical Thick-slicE MRI scans, bypassing both the costly SVR and the iterative non-rigid registration steps entirely, thereby substantially accelerating atlas construction. Extensive experiments demonstrate that INFANiTE outperforms existing baselines in subject consistency, reference fidelity, intrinsic quality and biological plausibility, even under challenging sparse-data settings. Additionally, INFANiTE reduces the end-to-end processing time (i.e., from raw scans to the final atlas) from days to hours compared to the traditional 3D volume-based pipeline (e.g., SyGN), facilitating large-scale population-level fetal brain analysis. Code: https://github.com/hu2274898/INFANiTE

URL PDF HTML 收藏
2604.25834 2026-07-15 cs.AI cs.IR 版本更新

Action-Aware Generative Sequence Modeling for Short Video Recommendation

面向动作的生成序列建模用于短视频推荐

Wenhao Li, Zihan Lin, Zhengxiao Guo, Jie Zhou, Shukai Liu, Yongqi Liu, Chuan Luo, Chaoyi Ma, Ruiming Tang, Han Li

机构 * Kuaishou Inc.(快手公司) Beihang University(北航)

AI总结 本文提出A2Gen模型,通过时间维度细化用户动作并生成序列,提升短视频推荐的准确性和用户参与度。

Comments We request the retraction of our paper due to discrepancies in the authorship information (https://dl.acm.org/doi/10.1145/3805712.3809728). To ensure accurate representation of contributions, we plan to revise the authorship list and resubmit. Thank you for your understanding

详情
AI中文摘要

随着互联网的快速发展,用户对内容消费平台推荐准确性的期望不断提高。然而,短视频通常包含多样化的片段,用户可能对所有片段持有不同的态度。传统二元分类推荐模型将视频视为单一整体实体,难以准确捕捉这种细微偏好。考虑到用户消费是时间过程,本文通过统计分析和动作模式研究,表明用户动作的时间点可以代表多样的意图。基于此,我们提出了一种新的建模范式:面向动作的生成序列网络(A2Gen),通过时间维度细化用户动作并将其链接成序列进行统一处理和预测。首先,我们引入了上下文感知注意力模块(CAM)来建模包含项目特定上下文特征的动作序列。在此基础上,我们开发了分层序列编码器(HSE)以从用户的历史动作中学习时间动作模式。最后,通过利用CAM,我们设计了一个动作序列生成模块:动作序列自回归生成器(AAG)。在快手数据集和天猫公开数据集上的大量离线实验表明,所提出模型的优越性。此外,通过在快手平台部署的大规模在线A/B测试,我们的模型在多任务预测中显著优于基线方法,通过利用序列信息,具体实现了用户观看时间增加0.34%、互动率增加8.1%、整体用户留存率(LifeTime-7)增加0.162%,成功在所有流量中部署,每天服务超过4亿用户。

英文摘要

With the rapid development of the Internet, users have increasingly higher expectations for the recommendation accuracy of online content consumption platforms. However, short videos often contain diverse segments, and users may not hold the same attitude toward all of them. Traditional binary-classification recommendation models, which treat a video as a single holistic entity, face limitations in accurately capturing such nuanced preferences. Considering that user consumption is a temporal process, this paper demonstrates that the timing of user actions can represent diverse intentions through statistical analysis and examination of action patterns. Based on this insight, we propose a novel modeling paradigm: Action-Aware Generative Sequence Network (A2Gen), which refines user actions along the temporal dimension and chains them into sequences for unified processing and prediction. First, we introduce the Context-aware Attention Module (CAM) to model action sequences enriched with item-specific contextual features. Building upon this, we develop the Hierarchical Sequence Encoder (HSE) to learn temporal action patterns from users' historical actions. Finally, through leveraging CAM, we design a module for action sequence generation: the Action-seq Autoregressive Generator (AAG). Extensive offline experiments on the Kuaishou's dataset and the Tmall public dataset demonstrate the superiority of our proposed model. Furthermore, through large-scale online A/B testing deployed on Kuaishou's platform, our model achieves significant improvements over baseline methods in multi-task prediction by leveraging sequential information. Specifically, it yields increases of 0.34% in user watch time, 8.1% in interaction rate, and 0.162% in overall user retention (LifeTime-7), leading to successful deployment across all traffic, serving over 400 million users every day.

URL PDF HTML 收藏
2607.11560 2026-07-14 cs.CV cs.AI 新提交

Technical Report on the CVPR 2026@AdvML Workshop Challenge

关于CVPR 2026@AdvML研讨会挑战赛的技术报告

Tianyuan Zhang, Zonglei Jing, Jiangfan Liu, Ligong Zhang, Ke Ma, Chengzhi Sun, Xiaohai Xu, Zhirui Zhang, Qianqian Xu, Qingming Huang, Hanyu Fang, Junhua Liu, Zheng Wang, Xiaoliang Liu, Yuanbo Li, Shuai Gui, Bin Wang, Menghe Zheng, Jing Nie, Hanyang Meng, Zeyang Zhang, Xiang Zhang, Yongxuan Zhu, Rui Ding, Hainan Li, Yongkang Zhang, Zhilei Zhu, Xianglong Kong, Jin Hu, Zonghao Ying, Yisong Xiao, Lei Chen, Haotong Qin, Jiakai Wang, Aishan Liu, Ruikai Li, Julia Karbing, Yinpeng Dong, Zhenfei Yin, Shao Jing, Xia Hu, Jingyi Xu, Juntao Dai, Xinyun Chen, Vishal M. Patel, Xianglong Liu, Dawn Song, Alan Yuille, Philip H. S. Torr, Dacheng Tao

机构 * Beihang University(北京航空航天大学) University of Chinese Academy of Sciences(中国科学院大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Tongji University(同济大学) iFLYTEK Co., Ltd.(科大讯飞股份有限公司) Anhui Laboratory for Safe Artificial Intelligence in the Yangtze River Delta(长三角安全人工智能安徽实验室) Wenzhou Business College(温州商学院) Jiangnan University(江南大学) Guangzhou City University of Technology(广州理工学院) Inceptio Technology(智元机器) Institute of Dataspace(数据空间研究所) Zhongguancun Laboratory(中关村实验室) Tsinghua University(清华大学) ETH Zürich(苏黎世联邦理工学院) University of Oxford(牛津大学) Shanghai AI Laboratory(上海人工智能实验室) BAAI(北京智源人工智能研究院) Meta(元公司) Johns Hopkins University(约翰·霍普金斯大学) University of California, Berkeley(加州大学伯克利分校) Nanyang Technological University(南洋理工大学)

AI总结 介绍CVPR 2026@AdvML研讨会针对自动驾驶VLAs的对抗性多模态攻击挑战赛,基于多视图视觉问答,参赛者要生成对抗图像和文本扰动。阐述任务设计等,研究领先提交作品发现后缀惩罚等模式,为多模态自动驾驶系统相关工作提供参考。

详情
AI中文摘要

视觉语言智能体(VLAs)越来越多地用于解释复杂驾驶场景并支持安全关键推理。本报告介绍了针对自动驾驶VLAs的对抗性多模态攻击的CVPR 2026@AdvML研讨会挑战赛。该挑战赛基于DriveLM风格的多视图视觉问答构建,用六个同步相机图像和一组结构化的驾驶相关问答对来表示每个场景。参与者生成对抗性图像和仅后缀的文本扰动,使模型响应偏离参考答案,同时保持图像保真度并限制文本成本。竞赛包括两个阶段,第二阶段增加了一个隐藏的黑盒模型来评估可迁移性。我们描述了任务设计、提交规则、评估协议和排行榜结果,然后研究了五份有技术报告的领先提交作品。在这些报告中出现了几个反复出现的模式:后缀惩罚有利于图像侧攻击;场景级、多视图优化比单独处理视图更有效;问答类型和图结构为分配攻击预算提供了有用的先验信息;特征空间目标可以提高黑盒迁移能力;相机图像中嵌入的排版内容暴露了驾驶VLAs中持续存在的漏洞。这些发现为未来多模态自动驾驶系统的鲁棒性评估和防御设计提供了实际参考。

英文摘要

Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built on DriveLM-style multi-view visual question answering, the challenge represents each scene with six synchronized camera images and a structured collection of driving-related question-answer pairs. Participants generate adversarial images and suffix-only textual perturbations that induce model responses to deviate from reference answers while preserving image fidelity and limiting textual cost. The competition comprises two phases, with Phase II adding a hidden black-box model to assess transferability. We describe the task design, submission rules, evaluation protocol, and leaderboard results, and then examine five leading submissions for which technical reports were available. Across these reports, several recurring patterns emerge: image-side attacks are favored by the suffix penalty; scene-level, multi-view optimization is more effective than treating views in isolation; QA types and graph structure provide useful priors for allocating attack budget; feature-space objectives can improve black-box transfer; and typographic content embedded in camera images exposes a persistent vulnerability in driving VLAs. These findings provide a practical reference for future robustness evaluation and defense design in multimodal autonomous-driving systems.

URL PDF HTML 收藏
2607.11529 2026-07-14 cs.CV 新提交

Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory

解析、搜索与确认:基于思维链推理和结构化空间记忆的免训练空中视觉与对话导航

Yu Qi, Hongyu Li, Shaofei Huang, Tianrui Hui, Yaxiong Wang, Lechao Cheng, Zhun Zhong, Si Liu, Meng Wang

机构 * School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机与信息工程学院) Jianghuai Advance Technology Center(江淮先进技术中心) Anhui Provincial Key Laboratory of Humanoid Robots(安徽省人形机器人重点实验室) School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) University of Macau(澳门大学)

AI总结 研究免训练环境下的空中视觉与对话导航任务,提出PSC - AVDN框架,结合解析 - 搜索 - 确认推理管道与结构化空间记忆,整合多种空间线索,在免训练设置中取得新的最优性能。

Comments 10 pages, 4 figures. Accepted to CVPR 2026

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 23859-23868

详情
AI中文摘要

本文针对资源高效的高空无人机免训练环境下的空中视觉与对话导航(AVDN)任务展开研究。应用大语言模型会因方向定位薄弱和缺乏明确空间线索导致导航不可靠。为此提出PSC - AVDN框架,将三阶段解析 - 搜索 - 确认推理管道与结构化空间记忆紧密结合。解析阶段用大语言模型转换指令,搜索思维链进行目标探索,确认思维链进行精细验证,结构化空间记忆整合多种空间线索。实验表明该框架在免训练环境中建立了新的最优性能。

英文摘要

In this paper, we tackle the Aerial Vision-and-Dialog Navigation (AVDN) task in the training-free setting for resource-efficient high-altitude UAV navigation.Naively applying MLLMs leads to unreliable navigation due to weak directional grounding and the lack of explicit spatial memory.To address these issues, we propose PSC-AVDN, a training-free framework that tightly couples a three-stage Parsing-Search-Confirmation reasoning pipeline with a Structured Spatial Memory (SSM).The parsing stage uses an LLM to convert ambiguous dialogue instructions into stable geometric directional and destination cues.A Search Chain-of-Thought (S-CoT) then performs stepwise target exploration under high-altitude observations, and a Confirmation Chain-of-Thought (C-CoT) conducts fine-grained verification around candidate regions to resolve visual ambiguity.Meanwhile, SSM integrates three complementary sources of spatial cues, including multi-scale visual observation, spatial visual memory, and structured geometric memory to provide global spatial context and long-horizon consistency.Extensive experiments on ANDH and ANDH-Full show that PSC-AVDN establishes new state-of-the-art performance in the training-free setting, matching or surpassing several finetuned methods.Code will be publicly available at: https://github.com/QY6616/PSC-AVDN

URL PDF HTML 收藏
2607.11374 2026-07-14 cs.LG 新提交

Surprisingly Simple and Effective Multi-Domain Graph Foundation Model through Graph-to-Table Alignment

通过图到表对齐实现惊人简单且有效的多域图基础模型

Chunyu Hu, Tianyin Liao, Ge Lan, Xingxuan Zhang, Jianxin Li, Peng Cui, Ziwei Zhang

机构 * Nankai University(南开大学) Beihang University(北京航空航天大学) Tsinghua University(清华大学)

AI总结 研究如何让表格基础模型有效捕获图结构信息,提出GTAlign框架,先预训练图编码器捕获图表示,通过社区引导持续预训练弥合差距,应用于目标域推理,实验表明该框架在节点和图分类上显著优于基线,提供简单无文本的图基础模型。

详情
AI中文摘要

图基础模型(GFMs)已成为跨不同图域学习可转移表示的有前途的范式。GFMs的最新进展主要由图神经网络和基于大语言模型(LLM)的方法主导,但这些方法在有限数据训练和对文本属性的严重依赖之间面临根本困境。表格基础模型(TFMs)提供了一种潜在替代方案,但如何让TFMs有效捕获图的结构信息仍未得到充分探索。关键挑战是学习一种图到表对齐机制。为解决此问题,我们提出了GTAlign,一种用于无文本图基础模型的惊人简单而有效的图到表对齐框架。具体而言,我们首先预训练一个图编码器,将不同的图映射到统一的潜在空间以捕获与域无关的图表示。为进一步弥合图拓扑与表格表示空间之间的差距,我们提出社区引导的持续预训练,使用从图社区派生的伪标签来构建少样本预测情节。最后,我们将图编码器应用于未见的目标域并进行上下文推理。在五个基准数据集上的广泛实验表明,GTAlign在节点和图分类方面均显著优于现有基线,提供了一个简单、有效且无文本的GFM模型。代码将在接受后发布。

英文摘要

Graph Foundation Models (GFMs) have emerged as a promising paradigm for learning transferable representations across diverse graph domains. Recent advancements in GFMs have been largely dominated by two paradigms: Graph Neural Network and Large Language Model (LLM) based methods. However, these methods often face a fundamental dilemma between training with limited data and a heavy reliance on textual attributes. Tabular foundation models (TFMs) offer a potential alternative, as node features and representations can be naturally organized in a tabular form. However, how to enable TFMs to effectively capture structural information of graphs remains largely unexplored. The key challenge is to learn a graph-to-table alignment mechanism that enables graph structural understanding for TFMs. To address this, we propose GTAlign, a surprisingly simple yet effective Graph-to-Table Alignment framework for text-free Graph Foundation Model. Specifically, we first pretrain a graph encoder that maps diverse graphs into a unified latent space to capture domain-agnostic graph representations. To further bridge the gap between graph topology and the tabular representation space, we propose community-guided continual pre-training, where pseudo-labels derived from graph community are used to construct few-shot prediction episodes. Lastly, we adapt the graph encoder for an unseen target domain and perform in-context inference. Extensive experiments on five benchmark datasets demonstrate that GTAlign significantly outperforms state-of-the-art baselines on both node and graph classification, offering a simple, effective, and text-free GFM model. Code will be released upon acceptance.

URL PDF HTML 收藏
2607.11340 2026-07-14 cs.RO 新提交

CR-Solver: GPU-Accelerated Kinematics Solver for Tendon-driven Continuum Robots

CR-Solver:用于腱驱动连续体机器人的GPU加速运动学求解器

Heqing Yang, Yang Yi, Linqing Zhong, Linjiang Huang, Si Liu

机构 * Beihang University(北京航空航天大学)

AI总结 针对连续体机器人运动规划问题,提出CR-Solver求解器,它基于优化框架统一逆运动学等,利用GPU加速并行优化,在三项任务验证中比传统CPU求解器显著加速,成功率超95%精度达毫米级,且用纯Python实现,为高性能运动规划提供基础。

Comments IROS 2026

详情
AI中文摘要

连续体机器人具有内在柔顺性、高灵活性和安全的物理交互性,可在受限和非结构化环境中导航与操作。尽管传感和控制方面有进展,但多数规划库基于刚体假设,缺乏适用于连续体机器人的快速实用工具。为此提出CR-Solver,一种用于腱驱动连续体机器人运动生成的两阶段、基于优化的求解器。该方法在单个约束非线性优化框架中统一了逆运动学、路径跟踪和轨迹规划。利用GPU加速并行优化,能快速、准确且考虑约束地给出解决方案。在三项任务上验证,比传统CPU求解器显著加速,成功率超95%且精度达毫米级。该求解器用纯Python实现降低了采用门槛,为连续体机器人高性能运动规划提供实用且可扩展基础。

英文摘要

Continuum robots provide intrinsic compliance, high dexterity, and safe physical interaction, enabling navigation and manipulation in confined and unstructured environments. Despite recent advances in sensing and control, heightening the need for precise motion generation, most widely used planning libraries are grounded in rigid-body assumptions, creating a critical gap for fast and practical tools for continuum robots. To address this, we present CR-Solver, a two-stage, optimization-based solver for the motion generation of tendon-driven continuum robots. Our method unifies inverse kinematics, path following, and trajectory planning within a single constrained nonlinear optimization framework. Leveraging GPU-accelerated parallel optimization, CR-Solver delivers fast, accurate, and constraint-aware solutions. We validate our approach on three tasks, demonstrating significant speedups over traditional CPU-based solvers while achieving a consistently high success rate above 95% and millimeter-level accuracy. The solver is implemented in pure Python, reducing the barrier to adoption and offering a practical, extensible foundation for continuum robots' high-performance motion planning.

URL PDF HTML 收藏
2607.11008 2026-07-14 cs.CV cs.AI 新提交

SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception

SynCLIP:用于鲁棒开放词汇密集感知的同义词一致语言-图像预训练

Mingjie Xie, Guangjun He, Dongli Xu, Youtian Lin, Hongjue Li, Pengming Feng, Jian Guan, Yue Deng

机构 * Beihang University(北京航空航天大学) State Key Laboratory of Space Information System and Integrated Application(空间信息系统与集成应用国家重点实验室) Nanjing University(南京大学) Harbin Engineering University(哈尔滨工程大学) Beijing Zhongguancun Academy(北京中关村科学城)

AI总结 研究针对开放词汇密集感知中同义词引起的定位不一致问题,提出SynCLIP框架,通过SSA和SAR模块增强注意力一致性与定位精度,构建SEViC支持预训练,实验证明该方法显著提升定位一致性并达领先性能。

Comments Accepted by CVPR 2026

详情
AI中文摘要

开放词汇密集感知(OVDP)旨在通过利用文本知识来定位训练期间未见过的物体。尽管基于CLIP的方法最近取得了显著进展,但存在一个关键限制:同义词引起的定位不一致,即语义等效的表达式会产生不同的空间注意力模式。这种不一致削弱了现有方法在实际OVDP应用中的鲁棒性和性能。为解决此问题,我们提出了SynCLIP,一个同义词一致语言-图像预训练框架,用于增强OVDP的同义词鲁棒定位。SynCLIP引入了语义一致空间注意力对齐(SSA)模块,通过最小化原始和同义词表达式的注意力图之间的差异来增强空间注意力一致性。此外,空间注意力细化(SAR)模块在对齐图中选择性地加强最语义相关的空间区域,以实现更精确和稳定的定位。为支持同义词一致预训练,我们还构建了一个同义词丰富视觉语料库(SEViC),用多个同义词和文本定义扩充每个类别。在多个基准上的广泛实验表明,SynCLIP在不同语言变体下显著提高了定位一致性,并在基于CLIP的OVDP方法中取得了领先性能。

英文摘要

Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remarkable progress of recent CLIP-based approaches, we identify a critical limitation: synonym-induced grounding inconsistency, where semantically equivalent expressions yield disparate spatial attention patterns. This inconsistency undermines the robustness and performance of existing methods in real-world OVDP applications. To address this issue, we propose SynCLIP, a Synonym-Coherent Language-Image Pretraining framework that enhances synonym-robust grounding for OVDP. SynCLIP introduces a Semantic-consistent Spatial Attention alignment (SSA) module to enhance spatial attention consistency by minimizing discrepancies between attention maps of original and synonymous expressions. Furthermore, a Spatial Attention Refinement (SAR) module selectively strengthens the most semantically relevant spatial regions within aligned maps for more precise and stable grounding. To support synonym-coherent pretraining, we also construct a Synonym-Enriched Visual Corpus (SEViC), which augments each category with multiple synonyms and textual definitions. Extensive experiments on multiple benchmarks demonstrate that SynCLIP substantially improves grounding consistency under diverse linguistic variants and achieves state-of-the-art performance among CLIP-based OVDP methods. Code is available at https://github.com/Justlovesmile/SynCLIP.

URL PDF HTML 收藏
2607.10805 2026-07-14 cs.CL cs.LG 新提交

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

诊断和缓解基于策略的自蒸馏中的思维崩溃

Keqin Peng, Chen Li, Yuanxin Ouyang, Yancheng Yuan, Liang Ding

机构 * Beihang University(北京航空航天大学) Hong Kong Polytechnic University(香港理工大学) Alibaba Group(阿里巴巴集团)

AI总结 研究基于策略的自蒸馏在复杂推理任务中性能下降问题,发现思维崩溃陷阱。提出自适应双视角OPSD,通过动态调节自蒸馏目标缓解该问题,在多模型规模和数据集上实验验证其有效性,提升绝对平均准确率4.1%。

详情
AI中文摘要

基于策略的自蒸馏(OPSD)已成为增强和对齐大语言模型(LLMs)的关键范式。然而,在复杂推理任务中,OPSD反而会降低下游性能。本文系统研究了这一问题,发现了一种严重的优化陷阱——思维崩溃,即模型原生中间推理行为急剧下降。通过基于熵的梯度掩码和令牌级目标分析,揭示了崩溃的触发机制。为解决此问题,提出了自适应双视角OPSD(AD - OPSD),通过不对称逐点散度门动态调节自蒸馏目标。实验表明,AD - OPSD在不同模型规模和数据集上比标准OPSD绝对平均准确率提高了4.1%,能缓解思维崩溃并能稳健推广。

英文摘要

On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxically degrades downstream performance. In this paper, we systematically investigate this pathology and identify a severe optimization trap we define as \textbf{Thinking Collapse} -- a sharp decline in the model's native intermediate reasoning behavior, measured by epistemic-token density (ET per 1k). Through entropy-based gradient masking and token-level target analysis, we show that this collapse is triggered by aggressive teacher gradients at high-student-entropy decision forks, where student epistemic tokens are frequently suppressed into teacher non-epistemic targets and are highly concentrated in high pointwise student-teacher divergence regions. To resolve this optimization pathology, we propose \textbf{Adaptive Dual-Perspective OPSD (AD-OPSD)}, a robust control framework that dynamically moderates the self-distillation objective. AD-OPSD selectively anchors high-suppression-risk sandboxed tokens to a reference prior derived from the frozen base model via an asymmetrical pointwise divergence gate, preserving native thinking capacity while retaining OPSD's error-correcting power. Extensive experiments across competitive mathematical benchmarks show that AD-OPSD improves over standard OPSD by up to \textbf{+4.1\%} absolute average accuracy across diverse model scales and datasets. Further analysis demonstrates that AD-OPSD mitigates thinking collapse and generalizes robustly to different post-training paradigms.

URL PDF HTML 收藏
2607.10789 2026-07-14 cs.AI 新提交

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging

成像101:在科学计算成像上对大语言模型编码智能体进行基准测试

Siyi Chen, Jiahe Ying, Yixuan Jia, Yuxuan Gu, Enze Ye, Weimin Bai, Zhijun Zeng, Shaochi Ren, Binhong Gao, Yubing Li, Tianhan Zhang, He Sun

机构 * College of Future Technology and the National Biomedical Imaging Center, Peking University(北京大学未来技术学院和国家生物医学成像中心) AI for Science Institute (AISI)(科学人工智能研究所) University of Michigan(密歇根大学) State Key Laboratory of Acoustics and Marine Information, Institute of Acoustics, Chinese Academy of Sciences(中国科学院声学研究所声场声信息国家重点实验室) University of Chinese Academy of Sciences(中国科学院大学) School of Astronautics, Beihang University(北京航空航天大学宇航学院) Key Laboratory of Spacecraft Design Optimization and Dynamic Simulation Technologies, Ministry of Education(教育部航天器设计优化与动态仿真技术重点实验室)

AI总结 研究针对计算成像构建正确重建管道费力的问题,引入含57个任务的Imaging-101基准及三个评估轨道,评估七个前沿大语言模型,发现应用编码智能体于计算成像存在系统性挑战,指出技能增强、领域专业化智能体是可靠成像辅助的途径。

详情
AI中文摘要

计算成像从间接的、有噪声的测量中恢复隐藏信号,是跨学科定量发现的基础,但构建正确的重建管道需要深厚的领域专业知识,即使对于领域科学家来说也很费力。我们引入了Imaging-101,这是一个包含57个经过专家验证的计算成像任务的基准,涵盖六个科学领域,每个任务都基于一篇同行评审论文,并规范为标准化的四阶段管道(预处理、正向物理建模、逆求解器和可视化)。三个评估轨道(规划、功能级单元测试和端到端重建)在整个管道中探测不同的智能体能力。评估七个前沿大语言模型发现,将编码智能体应用于计算成像存在系统性挑战,这些挑战超出了一般编码基准所暴露的问题,包括算法选择、物理惯例处理和管道集成。这些发现突出了具体的能力差距,并指出技能增强、领域专业化的智能体是实现可靠计算成像辅助的实际途径。

英文摘要

Computational imaging, which recovers hidden signals from indirect, noisy measurements, underpins quantitative discovery across scientific disciplines, yet building a correct reconstruction pipeline demands deep domain expertise and remains laborious even for domain scientists. We introduce Imaging-101, a benchmark of 57 expert-verified computational imaging tasks spanning six scientific domains, each grounded in a peer-reviewed paper and canonicalized into a standardized four-stage pipeline (preprocessing, forward physics modeling, inverse solver, and visualization) Three evaluation tracks (planning, function-level unit tests, and end-to-end reconstruction) probe distinct agent capabilities across the full pipeline. Evaluating seven frontier LLMs uncovers systematic challenges in applying coding agents to computational imaging that go beyond those exposed by general coding benchmarks, spanning algorithm selection, physical convention handling, and pipeline integration. These findings highlight concrete capability gaps and point toward skill-augmented, domain-specialized agents as a practical path to reliable computational imaging assistance.

URL PDF HTML 收藏
2607.08765 2026-07-14 cs.CV 版本更新

Enhancing In-context Panoramic Generation via Geometric-aware Pretraining

通过几何感知预训练增强上下文全景生成

Haoran Feng, Ruiyang Zhang, Longyi Zhang, Dizhe Zhang, Lu Qi

机构 * Insta360 Research(影石创新研究院) Tsinghua University(清华大学) Beihang University(北京航空航天大学) Wuhan University(武汉大学)

AI总结 该研究提出两阶段框架Canvas360用于上下文全景生成,结合几何感知预训练与微调。通过构建数据集及特定建模方法,增强文本到全景生成,经实验验证其能提高全景图像保真度,在相关指标上表现优异。

Comments Project page: https://zry000.github.io/Canvas360/ Github: https://github.com/Insta360-Research-Team/Canvas360

详情
AI中文摘要

在这项工作中,我们提出了Canvas360,这是一个用于上下文全景生成的两阶段框架,它将几何感知预训练与下游任务特定的微调相结合。为解决缺乏针对上下文全景任务的大规模、高质量训练数据的问题,我们提出了Canvas360Dataset,它包含100万个用于风格迁移、图像修复、图像扩展和编辑的高质量配对全景样本。在建模方面,Canvas360通过并行深度生成、速度循环填充和相似性损失正则化来增强文本到全景的生成。实验表明,Canvas360提高了全景图像保真度,在全景特定的FAED指标上表现出色,并在定量评估中取得了有竞争力或领先的结果。

英文摘要

In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream task-specific fine-tuning. To address the lack of large-scale, high-quality training data tailored to in-context panoramic tasks, we propose Canvas360Dataset, a collection of 1M high-quality paired panoramic samples for style transfer, inpainting, outpainting, and editing, enabling effective supervision across diverse in-context generation scenarios. On the modeling side, Canvas360 enhances text-to-panorama generation through parallel depth generation, velocity circular padding, and similarity loss regularization, enabling the model to learn geometry-aware representations, capture object distortion details, and improve geometric consistency and global coherence. Furthermore, empowered by strong panoramic priors, Canvas360 enables a unified in-context panoramic generation framework that supports diverse downstream tasks via token-level concatenation, surpassing prior methods in both task coverage and modeling flexibility. Extensive experiments show that Canvas360 improves panoramic image fidelity, achieving particularly strong performance on the panorama-specific FAED metric and competitive or leading results across the reported quantitative evaluations. More information can be found on our project page: https://zry000.github.io/Canvas360/

URL PDF HTML 收藏
2605.14712 2026-07-14 cs.RO cs.AI cs.CL cs.CV 版本更新

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

IntentVLA:用于有歧义机器人操作的短周期意图建模

Shijie Lian, Bin Yu, Xiaopeng Lin, Zhaolong Shen, Laurence Tianruo Yang, Yurun Jin, Haishan Liu, Changti Wu, Hang Yuan, Cong Huang, Kai Chen

机构 * HUST(华中科技大学) ZGCA(中钢集团人工智能研究院) ZGCI(中钢智能科技有限公司) HIT(哈尔滨工业大学) HKUST(GZ)(香港科技大学(广州)) BUAA(北京航空航天大学) ZZU(浙江工业大学) ECNU(华东师范大学) USTC(中国科学技术大学) DeepCybo

AI总结 IntentVLA通过编码近期视觉观测为紧凑的短周期意图表示,提升机器人操作的执行稳定性,并在多个基准测试中优于现有VLA基线。

Comments Code can be found in https://github.com/ZGC-EmbodyAI/IntentVLA

详情
AI中文摘要

机器人模仿数据通常是多模态的:相似的视觉-语言观测可能被不同的动作片段所跟随,因为人类演示者具有不同的短周期意图、任务阶段或最近的上下文。现有的基于帧的VLA策略仅根据当前观测和指令推断每个片段,因此在部分可观测性下可能在相邻重规划步骤中重新采样不同的意图,导致片段间冲突和不稳定执行。我们引入了IntentVLA,一种基于历史的VLA框架,将最近的视觉观测编码为紧凑的短周期意图表示,并用它来条件化片段生成。我们进一步引入了AliasBench,一个包含12个任务的歧义意识基准,在RoboTwin2上配有匹配的训练数据和评估环境,该环境隔离了短周期观测歧义。在AliasBench、SimplerEnv、LIBERO和RoboCasa中,IntentVLA提高了轨迹稳定性,并在多个基准测试中优于强大的VLA基线。

英文摘要

Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter-chunk conflict and unstable execution. We introduce IntentVLA, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12-task ambiguity-aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short-horizon observation aliasing. Across AliasBench, SimplerEnv, LIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines

URL PDF HTML 收藏
2511.02776 2026-07-14 cs.RO 版本更新

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

XR-1:通过学习统一的视觉-运动表示实现多功能的视觉-语言-动作模型

Shichao Fan, Kun Wu, Zhengping Che, Xinhua Wang, Di Wu, Fei Liao, Ning Liu, Yixue Zhang, Zhen Zhao, Zhiyuan Xu, Meng Li, Qingjie Liu, Shanghang Zhang, Min Wan, Jian Tang

机构 * Beijing Innovation Center of Humanoid Robotics, Beijing, China(北京人形机器人创新中心,北京,中国) School of Mechanical Engineering and Automation, Beihang University, Beijing, China(北京航空航天大学机械工程及自动化学院,北京,中国) State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,SCSE,北京航空航天大学,北京,中国) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(多媒体信息处理国家重点实验室,计算机科学学院,北京大学,北京,中国)

AI总结 XR-1通过学习统一的视觉-运动表示,解决视觉-语言-动作模型在低级动作生成和跨数据源领域差距的挑战,提出三阶段训练方法并验证了其在多种机器人和任务上的优越性能。

Comments Accepted to ICML2026 as Oral

详情
AI中文摘要

XR-1通过学习统一的视觉-运动表示,解决视觉-语言-动作模型在低级动作生成和跨数据源领域差距的挑战,提出三阶段训练方法并验证了其在多种机器人和任务上的优越性能。

英文摘要

Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still face two fundamental challenges: (i) producing precise low-level actions from high-dimensional observations, (ii) bridging domain gaps across heterogeneous data sources, including diverse robot embodiments and human demonstrations. Existing methods often encode latent variables from either visual dynamics or robotic actions to guide policy learning, but they fail to fully exploit the complementary multi-modal knowledge present in large-scale, heterogeneous datasets. In this work, we present X Robotic Model 1 (XR-1), a novel framework for versatile and scalable VLA learning across diverse robots, tasks, and environments. XR-1 introduces the \emph{Unified Vision-Motion Codes (UVMC)}, a discrete latent representation learned via a dual-branch VQ-VAE that jointly encodes visual dynamics and robotic motion. UVMC addresses these challenges by (i) serving as an intermediate representation between the observations and actions, and (ii) aligning multimodal dynamic information from heterogeneous data sources to capture complementary knowledge. To effectively exploit UVMC, we propose a three-stage training paradigm: (i) self-supervised UVMC learning, (ii) UVMC-guided pretraining on large-scale cross-embodiment robotic datasets, and (iii) task-specific post-training. We validate XR-1 through extensive real-world experiments with more than 14,000 rollouts on six different robot embodiments, spanning over 120 diverse manipulation tasks. XR-1 consistently outperforms state-of-the-art baselines such as $π_{0.5}$, $π_0$, RDT, UniVLA, and GR00T-N1.5 while demonstrating strong generalization to novel objects, background variations, distractors, and illumination changes. Our project is at https://xr-1-vla.github.io/.

URL PDF HTML 收藏