arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Nanjing University(南京大学)

至 收录 1217
2607.17900 2026-07-21 cs.SD 新提交

Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer

利用TTS:借助利用层实现上下文感知的富有表现力的语音合成

Shengfan Shen, Di Wu, Xingchen Song, Dinghao Zhou, Pengyu Cheng, Sixiang Lyu, Jian Luan, Shuai Wang

机构 * Xiaomi Inc.(小米公司) Nanjing University(南京大学)

AI总结 研究针对语音助手富有表现力的语音合成中灵活风格控制的需求,提出Harness TTS轻量级控制层,通过封闭集提示 - 工具路由实现风格控制,实验证明其在路由和合成任务中表现出色,为语音助手表达控制提供实用方案。

详情
AI中文摘要

语音助手的富有表现力的语音合成需要灵活的风格控制,以适应明确请求和更广泛的交互上下文。我们提出了Harness TTS,这是一个轻量级控制层,围绕TTS引擎以外部化并管理其表达行为。它将风格控制重新制定为封闭集提示 - 工具路由:离线时,使用结构化元数据构建风格提示工具的紧凑注册表;在线时,一个大语言模型规划器根据优先级感知观察模式选择合适的工具,TTS执行器使用相应的提示音频合成语音。我们在路由和合成任务上评估了Harness TTS。在路由方面,Qwen3 - 4B在显式、隐式和冲突子集上的Top - 1准确率分别为74.3%、43.0%和64.6%。对于合成,在CosyVoice3和VoxCPM2上的实验表明,Harness TTS优于仅指令控制,实现了更高的指令跟随胜率(在CosyVoice3上优势为23.1 - 35.6分,在VoxCPM2上为13.8 - 20.0分),并将UTMOSv2分数提高了0.11 - 0.38。此外,4B规划器在标准模式下不到50毫秒就能给出第一个工具推荐,对实时交互引入的延迟可忽略不计。这些结果表明,为TTS引擎配备专用的利用层为语音助手表达控制提供了一个实用、可审计且上下文感知的解决方案。

英文摘要

Expressive speech synthesis for voice assistants requires flexible style control that adapts to explicit requests and broader interaction context. We propose Harness TTS, a lightweight control layer that wraps around a TTS engine to externalize and govern its expressive behavior. It reformulates style control as closed-set prompt-tool routing: offline, a compact registry of stylistic prompt tools is constructed with structured metadata; online, an LLM planner selects the appropriate tool based on a priority-aware observation schema, and the TTS executor synthesizes speech using the corresponding prompt audio. We evaluate Harness TTS on both routing and synthesis tasks. In routing, Qwen3-4B achieves Top-1 accuracies of 74.3%, 43.0%, and 64.6% on explicit, implicit, and conflict subsets. For synthesis, experiments on CosyVoice3 and VoxCPM2 show that Harness TTS outperforms instruction-only control, achieving higher instruction-following win rates (margins of 23.1-35.6 points on CosyVoice3 and 13.8-20.0 points on VoxCPM2) and improving UTMOSv2 scores by 0.11-0.38. Moreover, the 4B planner delivers its first tool recommendation in under 50 ms in standard mode, introducing negligible latency for real-time interaction. These results demonstrate that equipping TTS engines with a dedicated Harness layer offers a practical, auditable, and context-aware solution for voice assistant expression control.

URL PDF HTML 收藏
2607.17646 2026-07-21 cs.RO 新提交

Configuration-Induced Passive Self-Rotation for Perception-Enhanced Autonomous Flight

用于感知增强自主飞行的构型诱导被动自旋转

Xurui Liu, Miao Wang, Tianyu He, Linrui Yang, Xiaobin Zhou

机构 * School of Robotics and Automation, Nanjing University(南京大学机器人与自动化学院)

AI总结 研究受限杂乱环境中自主飞行因传感器视野受限的问题,提出构型诱导被动自旋转三旋翼飞行器,利用后臂构型参数调节工作点,开发分层自主框架,经大量实验验证该方法对感知增强自主飞行有效。

Comments 9 pages, 7 figures, 4 tables

详情
AI中文摘要

在受限且杂乱的环境中进行自主飞行,机载传感器有限的视野从根本上限制了其发展。被动自旋转可在无需额外传感器的情况下扩大传感覆盖范围,但在扫描视野刷新率和飞行性能之间存在权衡。本文提出了一种用于感知增强自主飞行的构型诱导被动自旋转三旋翼飞行器。首先,利用后臂构型参数调节被动自旋转工作点,提供一种机身级机制来平衡扫描视野刷新率和飞行性能。其次,开发了一个集成规划与控制的分层自主框架,以在持续被动自旋转下实现敏捷且稳健的自主飞行。对于基于航点的检查,进一步使用引导点重新规划来提高任务级覆盖范围。大量实际实验,包括高速轨迹跟踪、抗干扰测试以及在典型杂乱环境中的自主导航,证明了所提方法对感知增强自主飞行具有有效性。

英文摘要

Autonomous flight in confined and cluttered environments is fundamentally limited by the restricted field of view (FoV) of onboard sensors. Passive self-rotation expands sensing coverage without additional sensors but introduces a tradeoff between swept-FoV refresh rate and flight performance. This letter presents a configuration-induced passively self-rotating tricopter for perception-enhanced autonomous flight. Firstly, the rear-arm configuration parameter is exploited to regulate the passive self-rotation operating point, providing an airframe-level mechanism for balancing swept-FoV refresh rate and flight performance. Secondly, a hierarchical autonomy framework integrating planning and control is developed to enable agile and robust autonomous flight under continuous passive self-rotation. For waypoint-based inspection, guide-point replanning is further used to improve task-level coverage. Extensive real-world experiments, including high-speed trajectory tracking, disturbance-rejection tests, and autonomous navigation in representative cluttered environments, demonstrate the effectiveness of the proposed approach for perception-enhanced autonomous flight.

URL PDF HTML 收藏
2607.17599 2026-07-21 cs.CV 新提交

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

ConsiSpace:学习几何一致性对视频空间推理很重要

Ting Huang, Zhenyu Zhang, Wenyuan Huang, Jian Yang, Hao Tang

机构 * School of Intelligent Science and Technology, Nanjing University(南京大学智能科学与技术学院) School of Computer Science, Peking University(北京大学计算机科学学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

AI总结 针对视频空间推理中现有模型难以聚合一致空间证据的问题,提出ConsiSpace框架,通过构建几何一致内存及采用统一一致性自监督强化学习,在三个空间推理基准测试上取得成绩提升,平均得分比最强基线高12.6分。

Comments ECCV 2026

详情
AI中文摘要

视频空间推理对导航感知和长视频问答至关重要,模型需在变化视角下推断长距离空间关系。但现有多模态大语言模型以语义为中心,难以可靠聚合冗余视频观测中的一致空间证据。为此提出ConsiSpace,一个几何一致性感知框架,将空间一致性转化为证据组织原则和明确的后SFT学习信号。构建几何一致内存,利用高效组织策略保存空间证据,通过统一一致性自监督强化学习提升跨视图稳定性。在三个基准测试上实验显示成绩提升,比最强基线平均得分提高12.6分。

英文摘要

Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal. We build a geometry-consistent memory (GCM) including implicit evidence tokens and explicit geometric cues, and leverage efficient organization strategies to compactly preserve task-related spatial evidence. Furthermore, we utilize unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability, with answer-, metric-, and topology-consistency rewards. Extensive experiments on three spatial-reasoning benchmarks, VSI-Bench, OSI-Bench, and MMSI-Video-Bench, show consistent gains, improving the average score by 12.6 points over the strongest baselines.

URL PDF HTML 收藏
2607.17585 2026-07-21 cs.CV 新提交

Pixel-Space Diffusion Transformers

像素空间扩散变换器

Renye Yan, Jikang Cheng, You Wu, Ling Liang, Wei Peng, Athanasios V. Vasilakos, Qingyu Zhao, Yu Zhang, Ehsan Adeli, Kilian M. Pohl, Guoying Zhao

机构 * Peking University(北京大学) Nanjing University(南京大学) Stanford University(斯坦福大学) Cornell University(康奈尔大学) University of Oulu(奥卢大学)

AI总结 本文探讨像素空间扩散变换器,针对潜在扩散模型的局限,研究直接对原始像素建模的像素空间扩散方法,介绍其在高维建模中的挑战与多模态建模优势,从多方面回顾pDiTs,总结方法、识别挑战并展望未来方向。

详情
AI中文摘要

潜在扩散模型(LDMs)通过在VAE压缩的潜在空间中去噪来实现高效的高分辨率图像合成。然而,固定视觉tokenizer会丢弃精细纹理和结构细节,且单独的表示和扩散训练会导致重建与生成目标不匹配。这些限制使人们重新关注像素空间扩散,它直接对原始像素建模,消除VAE瓶颈并支持端到端优化。这虽更符合高保真生成需求,但在高维建模中带来挑战。像素空间建模也为统一多模态系统提供了基础。本文从模型架构、连续生成机制和统一多模态建模角度回顾像素空间扩散变换器(pDiTs),总结方法、识别挑战并探讨未来方向。

英文摘要

Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual tokenizers can discard fine textures and structural details, while separate representation and diffusion training creates a mismatch between reconstruction and generation objectives. These limitations have renewed interest in pixel-space diffusion, which models raw pixels directly, removes the VAE bottleneck, and supports end-to-end optimization. This formulation better matches the demands of high-fidelity generation but introduces challenges in high-dimensional modeling, including noise scheduling, loss weighting, token efficiency, and scalable architecture design. Pixel-space modeling also offers a promising basis for unified multimodal systems: raw pixels, text, and task conditions can be represented in a shared token space and jointly processed by a single Transformer, narrowing the gap between visual understanding and generation. This paper reviews Pixel-Space Diffusion Transformers (pDiTs) from the perspectives of model architecture, continuous generative mechanisms, and unified multimodal modeling. We summarize representative methods, identify key technical challenges, and discuss future directions toward high-fidelity, end-to-end vision foundation models that integrate generation and understanding.

URL PDF HTML 收藏
2607.17476 2026-07-21 cs.RO 新提交

Disturbance-Aware Flight for Aerial Robots in Narrow Space

狭窄空间中无人机的干扰感知飞行

Lei Qiang, Tianyu He, Chenyang Sun, Xurui Liu, Miao Wang, Xiaobin Zhou

机构 * School of Robotics and Automation, Nanjing University(南京大学机器人与自动化学院)

AI总结 针对无人机在狭窄空间自主飞行因干扰和空间受限面临的挑战,提出干扰感知规划与控制框架DAPCF,通过双环观测器估计干扰、引入干扰风险函数及设计MDNMPC,使四旋翼能穿越狭窄隧道,性能优于人类飞行员。

详情
AI中文摘要

由于强烈的空气动力学干扰和有限的飞行空间,无人机在狭窄空间的自主飞行仍然具有挑战性。现有方法主要在控制层面解决空气动力学干扰,而运动规划通常依赖几何约束和固定速度限制,在受限环境中导致保守或不安全行为。本文提出一种干扰感知规划与控制框架(DAPCF),将在线干扰估计集成到狭窄空间四旋翼飞行的规划控制回路中。首先,双环观测器基于里程计和电机速度测量实时估计六自由度干扰力和扭矩。然后,引入干扰风险函数,根据干扰估计自适应调节规划器的参考速度,在干扰超过阈值时降低速度,在低干扰条件下恢复速度。最后,设计基于电机动力学的带干扰补偿的非线性模型预测控制器(MDNMPC),以确保在扰动条件下的鲁棒轨迹跟踪。实验表明,对角线长度为0.39米的四旋翼可以穿越窄至0.6米的直、斜坡和弯曲隧道,在成功率和飞行效率方面均优于人类飞行员。

英文摘要

Autonomous flight of aerial robots in narrow space remains challenging due to strong aerodynamic disturbances and limited flying space. Existing approaches mainly address aerodynamic disturbances at the control level, while motion planning typically relies on geometric constraints and fixed speed limits, leading to conservative or unsafe behaviors in confined environments. This paper presents a disturbance-aware planning and control framework (DAPCF) that integrates online disturbance estimation into the planning-control loop for quadrotor flight in narrow space. First, the dual-loop observers estimate 6-degree-of-freedom disturbance forces and torques in real time based on odometry and motor speed measurements. Then, a disturbance risk function is introduced that adaptively modulates the reference speed of the planner based on disturbance estimation, reducing velocity when disturbances exceed a threshold and restoring it under low-disturbance conditions. Finally, a motor-dynamics-based nonlinear model predictive controller (MDNMPC) with disturbance compensation is designed to ensure robust trajectory tracking under perturbed conditions. Experiments demonstrate that a quadrotor with a diagonal length of 0.39~m can traverse straight, sloped, and curved tunnels as narrow as 0.6~m, outperforming human pilots in both success rate and flight efficiency.

URL PDF HTML 收藏
2607.17423 2026-07-21 cs.CV 新提交

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

TimeLens2:使用多模态大语言模型进行通用视频时间定位

Yuhan Zhu, Changlian Ma, Xiangyu Zeng, Xinhao Li, Zhiqiu Zhang, Songze Li, Jun Zhang, Tianxiang Jiang, Yuandong Yang, Ziang Yan, Zikang Wang, Xinyu Chen, Haoran Chen, Shaowei Zhang, Limin Wang

机构 * Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) University of Science and Technology of China(中国科学技术大学) Fudan University(复旦大学)

AI总结 研究通用视频时间定位问题,TimeLens2将时间证据视为区间集,通过多种策略构建多跨度监督,其时间瓦瑟斯坦奖励等方法提供反馈,在多个基准测试中表现出色,不同变体均有性能提升。

Comments Technical Report

详情
AI中文摘要

视频多模态大语言模型能描述视频中发生的事情,但很少能确定支持证据出现的时间。我们研究通用视频时间定位,即一个模型预测跨视频长度、领域、查询形式和视角的可变基数证据区间集。现有训练策略与该集值任务不匹配。TimeLens2在整个监督和优化过程中将时间证据视为区间集。TimeLens2-93K通过字幕衍生提案、独立定位、跨智能体共识、语义验证和边界细化构建可靠的多跨度监督。我们的时间瓦瑟斯坦奖励计算合并区间支持上均匀分布之间精确的一维\(W_1\),在不等基数和等效碎片化情况下提供密集、无匹配反馈;时间交并比用精确重叠反馈进行补充。在七个基准测试中,TimeLens2-2B在每个基准上均优于所有大小匹配的基线,4B和8B变体达到了当前最优性能。2B、4B和8B变体分别比其Qwen3-VL主干提高了14.2、13.0和18.1的平均交并比分数。

英文摘要

Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints. Existing training strategies are misaligned with this set-valued task: long-video labels often rely on brittle one-pass annotation, while reinforcement-learning rewards either fail to distinguish non-overlapping predictions or require fragile segment matching. TimeLens2 treats temporal evidence as an interval set throughout supervision and optimization. TimeLens2-93K constructs reliable multi-span supervision through caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. Our temporal Wasserstein reward computes exact one-dimensional \(W_1\) between uniform distributions over merged interval supports, providing dense, matching-free feedback under unequal cardinalities and equivalent fragmentation; temporal IoU complements it with precise-overlap feedback. Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters. The 2B, 4B, and 8B variants improve over their Qwen3-VL backbones by 14.2, 13.0, and 18.1 mIoU points, respectively.

URL PDF HTML 收藏
2607.14661 2026-07-21 cs.AI 版本更新

SmartRAG: Native Graph-Based RAG for Mobile Device

SmartRAG:用于移动设备的基于原生图的RAG

Zhihan Jiang, Meng Li, Shenghao Liu, Keran Li, Ruiben Zhou, Wei Wang, Xianjun Deng, Shuai Wang, Haipeng Dai

机构 * Nanjing University(南京大学) HUST(华中科技大学) Southeast University(东南大学)

AI总结 研究针对移动设备部署大语言模型的问题,提出SmartRAG框架,围绕四个模块组织智能助手,核心是可持续学习的EvoNER,知识存于MRGraph,通过混合管道检索,实验表明其在多跳推理性能上有竞争力且能在普通智能手机上运行。

Comments 14 pages, 4 figures

详情
AI中文摘要

在移动设备上部署大语言模型作为个人助手需要隐私、低延迟和离线可用性,但巨型模型的计算成本与严格的边缘硬件预算相冲突。我们认为仅靠模型压缩无法解决这种矛盾,需要将设备上的智能分解为互补的功能角色。我们提出了SmartRAG,这是一个完全在设备上的框架,围绕四个协调模块——感知、记忆、聚焦和思考来组织智能助手。SmartRAG的核心是EvoNER,它是一个可持续学习的命名实体识别器,通过教师提炼更新逐步扩展其标签库。提取的知识存储在MRGraph中,并在查询时通过结合图遍历、词汇匹配和密集语义搜索的混合管道进行检索。仅在高价值语义操作时调用设备上的大语言模型,以限制推理成本。在四个问答基准测试上的实验表明,具有量化的17亿参数主干的SmartRAG实现了与高达18倍大的模型具有竞争力的多跳推理性能,同时完全在普通智能手机上运行,内存和延迟在实际范围内。

英文摘要

Deploying large language models (LLMs) as personal assistants on mobile devices demands privacy, low latency, and offline availability, yet the computational cost of giant models clashes with strict edge-hardware budgets. We argue that this tension cannot be resolved by model compression alone; it requires decomposing on-device intelligence into complementary functional roles. We present SmartRAG, a fully on-device framework that organizes an intelligent assistant around four coordinated modules -- Perception, Memory, Focus, and Thinking. At the core of SmartRAG is EvoNER, a continually learnable named-entity recognizer that incrementally expands its label inventory through teacher-distilled updates, enabling the system to absorb previously unseen entity types without retraining the backbone LLM. Extracted knowledge is stored in MRGraph, a three-layer provenance-preserving knowledge graph, and retrieved at query time through a hybrid pipeline combining graph traversal, lexical matching, and dense semantic search. The on-device LLM is invoked only for high-value semantic operations -- labeling, planning, and answer synthesis -- keeping inference costs bounded. Experiments on four QA benchmarks (TriviaQA, Natural Questions, HotpotQA, MultiHopQA) show that SmartRAG with a quantized 1.7B-parameter backbone achieves multi-hop reasoning performance competitive with models up to 18$\times$ larger, while running entirely on commodity smartphones within practical memory and latency envelopes.

URL PDF HTML 收藏
2607.07304 2026-07-21 cs.LG 版本更新

Nonlinear Bandit

非线性带式赌博机

Tianshuo Zheng, Ting Wu, Zhi-Hua Zhou, Keqin Liu

机构 * School of Mathematics, Nanjing University(南京大学数学系) School of Artificial Intelligence, Nanjing University, National Key Laboratory for Novel Software Technology(南京大学人工智能学院,计算机软件新技术国家重点实验室) School of Mathematics and Physics, Xi’an Jiaotong-Liverpool University(西交利物浦大学数理学院)

AI总结 研究重尾噪声下广义线性带式赌博机问题,基于在线镜像下降方法提出算法EHM,实现几乎最优遗憾值且无需常用参数,还研究了上下文特征分段常数时的情况及非线性带式赌博机特殊情况,给出相应算法并证明遗憾上界。

详情
AI中文摘要

本文首先研究重尾噪声下的广义线性带式赌博机(GLB)问题。重尾分布特征在个性化推荐、金融市场和医疗等实际应用中广泛存在。基于在线镜像下降(OMD)方法,我们提出算法EHM,它扩展了自适应Huber损失方法,实现几乎最优遗憾值$\widetilde{\mathcal{O}}(T^{\frac{1}{1+\epsilon}})$,且无需常用参数。接着研究上下文特征分段常数时的GLB问题,得到PGLB - EHM算法,遗憾上界阶不变。还深入研究非线性带式赌博机特殊情况,给出NB - EHM算法,最终利用仿射提升方法表明一般NB问题可用NB - EHM实现次线性遗憾界。

英文摘要

In this paper we first study the problem of generalized linear bandit (GLB) under heavy-tailed noise. The characteristics of heavy-tailed distributions are widely observed in real-world applications such as personalized recommendation, financial markets, and medical treatments. Based on the online mirror descent (OMD) method, we propose an algorithm EHM that extends the adaptive Huber loss method (Wang et al., 2025) with one-pass update ($\mathcal{O}(1)$ computational complexity with respect to current round $t$ and the time horizon $T$), which simultaneously achieves an almost optimal regret of $\widetilde{\mathcal{O}}(T^{\frac{1}{1+ε}})$ where $T$ is the time horizon. In addition, by utilizing a special property of some link function (Sawarni et al., 2025), our algorithm eliminates the need to know a commonly used parameter. Next, we study the GLB problem under the case when contextual characteristic becomes piecewise constant, and we slightly revised former algorithm to obtain the PGLB-EHM algorithm. After theoretical analysis, we prove that the regret upper bound order stays the same. Furthermore, we look deeper into a special case of nonlinear bandit (NB) and present the NB-EHM algorithm with bisection method and special restriction. Eventually we utilize the affine lifting approach and show that the general NB problem can be applied with NB-EHM to achieve a sublinear regret bound.

URL PDF HTML 收藏
2601.03043 2026-07-21 cs.CL cs.AI cs.LG 版本更新

Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage

Lil: 在长解码阶段应用后训练稀疏注意力算法时,少即是少

Junhao Hu, Fangze Li, Mingtao Xu, Feifan Meng, Shiju Zhao, Tiancheng Hu, Ting Peng, Anmin Liu, Wenrui Huang, Chenxu Liu, Ziyue Hua, Tao Xie

机构 * SCS, Peking University, Beijing, China(北京大学信息科学与技术学院,北京,中国) Key Lab of HCST (PKU), MOE, Beijing, China(高等教育出版社HCST重点实验室(PKU),北京,中国) State Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室,中国) Tencent, Shenzhen, China(腾讯,深圳,中国) Beijing Tongming Lake Information Technology Application Innovation Center, Beijing, China(北京 Tongming Lake 信息技术应用创新中心,北京,中国)

AI总结 本文研究了在长解码阶段应用稀疏注意力算法时,信息丢失导致序列变长的问题,提出早停算法减少token消耗并降低精度损失。

详情
AI中文摘要

大型语言模型(LLMs)在广泛复杂任务中表现出强大的能力,并且正在大规模部署,这对推理效率提出了显著要求。先前的工作通常将推理分解为prefill和decode阶段,其中decode阶段主导总延迟。为了减少解码阶段的时间和内存复杂度,一系列工作引入了稀疏注意力算法。在本文中,我们通过实证和理论证明,稀疏注意力可能反常地增加端到端复杂度:信息丢失往往导致显著更长的序列,这种现象我们称为“Less is Less”(Lil)。为缓解Lil问题,我们提出了一种早停算法,该算法检测稀疏解码过程中信息损失超过信息增益的阈值。我们的早停算法在推理密集型基准上将token消耗减少了高达90%,同时精度损失低于2%。

英文摘要

Large language models (LLMs) demonstrate strong capabilities across a wide range of complex tasks and are increasingly deployed at scale, placing significant demands on inference efficiency. Prior work typically decomposes inference into prefill and decode stages, with the decode stage dominating total latency. To reduce time and memory complexity in the decode stage, a line of work introduces sparse-attention algorithms. In this paper, we show, both empirically and theoretically, that sparse attention can paradoxically increase end-to-end complexity: information loss often induces significantly longer sequences, a phenomenon we term ``Less is Less'' (Lil). To mitigate the Lil problem, we propose an early-stopping algorithm that detects the threshold where information loss exceeds information gain during sparse decoding. Our early-stopping algorithm reduces token consumption by up to 90% with a marginal accuracy degradation of less than 2% across reasoning-intensive benchmarks.

URL PDF HTML 收藏
2510.23472 2026-07-21 cs.LG cs.AI cs.AR cs.NE 版本更新

BBOPlace-Bench: Benchmarking Black-Box Optimization for Chip Placement

BBOPlace-Bench:用于芯片布局的黑盒优化基准测试

Ke Xue, Ruo-Tong Chen, Rong-Xi Tan, Xi Lin, Yunqi Shi, Siyuan Xu, Mingxuan Yuan, Chao Qian

机构 * National Key Laboratory for Novel Software Technology(国家新型软件技术重点实验室) School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

AI总结 针对芯片布局中黑盒优化缺乏统一基准的问题,提出BBOPlace-Bench基准,整合问题表述、芯片案例和算法家族,提供模块化框架,通过共享协议系统评估算法性能,助力开发高效方案并拓宽应用场景。

Comments IEEE TEvC

详情
AI中文摘要

芯片布局是现代芯片设计的关键阶段,黑盒优化(BBO)已应用数十年。早期受限于问题表述和算法设计,效率等不如主流分析方法。虽近期BBO有进展,但缺乏统一基准。为此提出BBOPlace-Bench,首个针对芯片布局BBO算法评估与开发的基准。它整合三种问题表述,提供模块化框架,汇总芯片案例并标准化格式,集成算法家族并系统评估性能。在共享协议下,部分BBO配置具竞争力,既助力开发高效BBO驱动的芯片布局方案,又拓宽BBO社区急需的实际应用场景。

英文摘要

Chip placement is a vital stage in modern chip design, and black-box optimization (BBO) has been applied to it for decades. Early BBO efforts, however, were limited by immature problem formulations and inefficient algorithm designs, leading to worse efficiency, quality, and scalability than mainstream analytical methods. Recent advances in BBO have shown strong potential, but a unified, BBO-specific benchmark for thoroughly assessing various problem formulations and BBO algorithms is lacking. To fill this gap, we propose BBOPlace-Bench, the first benchmark tailored for evaluating and developing BBO algorithms for chip placement. It integrates three BBO problem formulations and offers a modular, flexible framework that enables users to seamlessly implement, test, and compare their own algorithms. It aggregates representative modern chip cases and standardizes their formats, providing uniform and comprehensive information to support BBO optimization. Moreover, it integrates representative BBO algorithm families, including simulated annealing, population-based search (including GA, CMA-ES, and PSO), and Bayesian optimization, and systematically evaluates their performance across different problem formulations using key chip-placement metrics. We position these experiments primarily as illustrative case studies under a shared evaluation protocol, including common benchmark instances, metric definitions, evaluation pipeline, and search budgets. Under this protocol, some BBO configurations (e.g., GA under the mask-guided optimization formulation) are competitive with representative analytical and reinforcement learning baselines. BBOPlace-Bench not only facilitates the development of efficient BBO-driven solutions for chip placement but also broadens the practical application scenarios urgently needed by the BBO community.

URL PDF HTML 收藏
2503.12562 2026-07-21 cs.CV 版本更新

History-Aware Transformation of ReID Features for Multiple Object Tracking

用于多目标跟踪的重识别特征的历史感知变换

Ruopeng Gao, Yuyao Wang, Chunxu Liu, Limin Wang

机构 * Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室)

AI总结 该研究针对多目标跟踪中ReID特征应用问题,提出历史感知特征变换方法,将轨迹历史特征作上下文,用定制FLD投影特征到特定空间,免训练提升特征判别力,证明特定上下文表示对MOT更优,推动ReID特征定制探索。

Comments Accepted by ECCV 2026. Without bells and whistles, achieving 80.8 HOTA on SportsMOT

详情
AI中文摘要

在多目标跟踪(MOT)中,重识别(ReID)特征被广泛用作对象关联的有力线索。然而,它们通常被简单地通过相似性度量统一应用于所有视频。我们认为这忽略了一个基本事实:MOT不是一般的检索问题,而是在单个视频中区分目标的特定上下文任务。为此,我们主张根据每个视频序列的特定上下文调整视觉特征以实现更好的适应。本文提出了一种历史感知特征变换方法,该方法根据每个视频的独特样本分布动态构建一个更具判别力的子空间。具体而言,我们将已建立轨迹的历史特征视为上下文,并采用定制的Fisher线性判别(FLD)将原始ReID特征投影到特定于序列的表示空间中。大量实验表明,我们的免训练方法显著增强了来自不同ReID主干的特征的判别力,从而在跟踪精度上取得了显著且一致的提升。我们的研究结果有力地证明,MOT本质上更倾向于特定上下文表示而非直接应用通用ReID特征。我们希望我们的工作能激励社区超越对ReID特征的简单应用,朝着更深入探索其针对MOT的有目的定制方向发展。我们的代码将发布。代码发布在这个https网址。

英文摘要

In Multiple Object Tracking (MOT), Re-identification (ReID) features are widely employed as a powerful cue for object association. However, they are often wielded as a one-size-fits-all hammer, applied uniformly across all videos through simple similarity metrics. We argue that this overlooks a fundamental truth: MOT is not a general retrieval problem, but a context-specific task of discriminating targets within a single video. To this end, we advocate for the adjustment of visual features based on the context specific to each video sequence for better adaptation. In this paper, we propose a history-aware feature transformation method that dynamically crafts a more discriminative subspace tailored to each video's unique sample distribution. Specifically, we treat the historical features of established trajectories as context and employ a tailored Fisher Linear Discriminant (FLD) to project the raw ReID features into a sequence-specific representation space. Extensive experiments demonstrate that our training-free method dramatically enhances the discriminative power of features from diverse ReID backbones, resulting in marked and consistent gains in tracking accuracy. Our findings provide compelling evidence that MOT inherently favors context-specific representation over the direct application of generic ReID features. We hope our work inspires the community to move beyond the naive application of ReID features and towards a deeper exploration of their purposeful customization for MOT. Our code will be released. The code is released at https://github.com/MCG-NJU/HATReID-MOT.

URL PDF HTML 收藏
2607.15593 2026-07-20 cs.DC cs.AI cs.NI 新提交

Scalable LLM Agent Tool Access in the Cloud

云中可扩展的大语言模型智能体工具访问

Mingxin Li, Enge Song, Yueshang Zuo, Xiaodong Liu, Rong Wen, Qiang Fu, Gianni Antichi, Jian He, Jing Tie, Zhou Shao, Xiaobo Xue, Xiong Xiao, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Zihao Fan, Haonan Li, Tian Pan, Xiaomin Wu, Yang Song, Xing Li, Biao Lyu, Meng Li, Haipeng Dai, Guihai Chen, Shunmin Zhu

机构 * Nanjing University(南京大学) Alibaba Cloud(阿里云) Fudan University(复旦大学) RMIT University(皇家墨尔本理工大学) Politecnico di Milano(米兰理工学院) Zhejiang University(浙江大学)

AI总结 研究大语言模型智能体在云环境下工具访问问题,提出云规模网关系统,通过打破直接连接模型等整合多种功能,实现高召回率、扩展工具访问数量、提高准确性并减少时间和令牌使用量,还分享了部署经验。

详情
AI中文摘要

大语言模型智能体越来越依赖工具调用与外部系统交互,模型上下文协议(MCP)成为事实上的接口。但在云规模下运行MCP困难重重。工具提供方存在遗留服务难通过MCP直接调用及协议开发带来兼容性成本问题。智能体方面,可访问工具数量受限于大语言模型上下文窗口和推理开销。本文提出云规模的网关系统,打破数据平面直接连接模型,卸载遗留服务集成,整合多种功能。混合检索召回率达98%,将智能体工具访问扩展到3000多个,提高工具选择准确性,减少工具选择时间和令牌使用量,每调用开销低,扩展时稳定。最后分享了生产中部署网关系统的经验教训。

英文摘要

LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de facto interface. Operating MCP at cloud scale, however, becomes difficult. On the tool provider side, legacy services are not directly callable through MCP; the rapid protocol development also creates ongoing compatibility cost. On the agent side, the number of accessible tool is limited by the LLM context window and inference overhead; mounting a large tool set increases token usage and inference latency and can reduce task success rate. Moreover, for stateful MCP backends with multiple replicas, preserving session affinity increases client-side complexity. We present a cloud-scale gateway system for MCP service. It breaks the direct-connect model on the data plane and offloads legacy service integration, consolidating incompatible MCP variants, access control, tool recommendation, and session-aware routing to the gateway. Hybrid retrieval sustains 98% Top-15 recall; it scales agent tool access to 3,000+ with high tool selection accuracy, and reduces tool selection time by $8.9\times$ and token usage by $23.8\times$, with low per-call overhead, stable under scale-out. Finally, we share the lessons learned from deploying the gateway system in production.

URL PDF HTML 收藏
2607.15198 2026-07-17 eess.AS cs.SD 新提交

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

SLT 2026真实TSE挑战:从对话录音中提取真实世界目标说话人

Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han, Zikai Liu, Xiaoyang Yu, Haoyu Li, Marc Delcroix, Kai Yu, Lei Xie, Ming Li, Haizhou Li

机构 * Nanjing University(南京大学) Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)) Brno University of Technology(布拉格技术大学) Northwestern Polytechnical University(西北工业大学) NTT, Inc.(NTT公司) Shanghai Jiao Tong University(上海交通大学)

AI总结 介绍SLT 2026的REAL-TSE挑战,从真实对话录音提取目标说话人,有在线和离线赛道,评估指标多样,描述了任务定义等多方面内容及经验教训。

Comments Overview paper of Real-TSE Challenge

详情
AI中文摘要

我们介绍了REAL-TSE挑战,这是IEEE SLT 2026关于从真实对话录音中提取目标说话人(TSE)的卫星挑战。给定多说话人混合语音和目标说话人的一个或多个注册话语,参与系统必须只恢复目标语音。与模拟朗读语音基准不同,REAL-TSE评估包含自然重叠、混响、噪声、信道失配和对话动态的普通话和英语录音。该挑战定义了两个互补赛道:用于低延迟流提取的在线赛道和用于全上下文处理的离线赛道。系统通过令牌错误率(TER)、说话人相似度(SpkSim)、DNSMOS和目标说话人活动F1进行评估。本文概述了任务定义、数据集、基线、评估协议、提交的系统、按条件的发现以及对未来真实世界TSE基准的经验教训。

英文摘要

We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.

URL PDF HTML 收藏
2607.14935 2026-07-17 cs.CV 新提交

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

VideoChat3:用于高效通用视频理解的全开放视频多模态语言模型

Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang

机构 * Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室) Nanyang Technological University(南洋理工大学) Peking University(北京大学)

AI总结 研究针对视频理解开源模型局限,提出VideoChat3。通过I3D-ViT等提升效率,利用可扩展视频数据合成管道生成训练数据集提升泛化性,以4B参数在多基准测试中超越同等或更多参数的开源模型,实现泛化与计算效率平衡。

详情
AI中文摘要

视频理解领域虽有进展,但当前开源模型存在局限。它们难以跨多种视频类型泛化,计算需求高限制了效率与可扩展性,且大多模型部分开放。为此引入全开放、高效且通用的以视频为中心的多模态语言模型VideoChat3。通过Inflated 3D Vision Transformer (I3D-ViT)等设计提升效率,开发可扩展视频数据合成管道生成三个训练数据集提升泛化性,实验表明其在泛化和计算效率上取得平衡,超越了同等或更多参数的开源模型。

英文摘要

Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-driven development. To address these issues, we introduce VideoChat3, a fully open, efficient, and generalist video-centric MLLM. VideoChat3 advances video understanding through two complementary designs. For efficiency, we introduce Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for Streaming Video Perception, which enables efficient spatiotemporal representation and reduces the cost of processing video inputs during training and inference. For effectiveness, we develop a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K, covering general, long-form, and streaming video scenarios, improving the model's generalization across domains. By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.

URL PDF HTML 收藏
2607.14660 2026-07-17 cs.CV 新提交

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

VIABench:一个从视障人士收集的用于视障辅助的综合视频基准测试

Yunfeng Liu, Yuandong Yang, Jiarui Han, Zhenpeng Huang, Yuqing Tang, Xiangyu Zeng, Gangshan Wu, Limin Wang

机构 * Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室)

AI总结 针对视障人士视觉信息获取难及多模态大语言模型在视障辅助中实用价值待探索的问题,引入VIABench视频基准测试,定义三个核心任务,提出严格测试流程,实验表明当前模型对视障人士支持不足,有望推动定制模型开发以改善其体验。

详情
AI中文摘要

视障人士因获取视觉信息有限而面临重大日常挑战。尽管多模态大语言模型在一般视觉和语言任务上取得了显著成果,但在实际视障辅助中的实用价值仍未充分探索。为填补这一空白,我们引入了VIABench,这是一个专门设计的综合视频基准测试,使用视障人士自己录制或分享的第一人称视频来评估多模态大语言模型在视障辅助场景中的表现。VIABench定义了三个核心任务,每个任务针对视觉辅助中的不同需求。主动提醒任务评估模型解释正在进行的视频内容并主动预测和口头描述即将到来的关键导航事件的能力;视觉问答任务评估模型回答用户关于视频中环境或物体问题的能力;视觉引导交互任务测试情境感知推理以完成用户与环境之间的有意交互。为确保进行稳健和公平的评估,我们提出了一个严格的基准测试流程,支持在线(实时)和离线设置。我们的实验表明,当前的多模态大语言模型仍难以对视障人士提供全面支持,尤其是在主动提醒任务中,该任务需要准确的预测和实时响应能力。我们希望VIABench能推动未来研究开发针对实际辅助的定制多模态大语言模型,最终改善视障人士的导航和交互体验。代码和数据将在这个https网址发布。

英文摘要

Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.

URL PDF HTML 收藏
2607.14651 2026-07-17 cs.CR cs.AI 新提交

MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents

MemPoison:揭示大语言模型智能体中持久内存威胁和结构盲点

Jifeng Gao, Kang Xia, Yi Zhang, Xiaobin Hong, Mingkai Lin, Xingshen Wei, Wenzhong Li, Sanglu Lu

机构 * State Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室) NARI Group Corporation/State Grid Electric Power Research Institute, China(国网电力研究院/国电电力集团)

AI总结 研究大语言模型智能体持久内存安全漏洞,提出MemPoison框架,含1227个案例,涵盖多种攻击类型等并评估多个模型系列。引入分类法,揭示防御边界及盲点,主张转向自适应、上下文敏感的内存防御策略。

详情
AI中文摘要

持久外部内存增强了智能体的连续性,但引入了持久的安全漏洞:对抗性内容可通过标准交互通道注入,跨轮保留并扭曲下游行为。为应对这一挑战,我们提出MemPoison,一个全面的基准测试和分析框架,包含1227个经过人工验证的案例,涵盖四种攻击类型、三种注入通道和三种代表性内存基板,在七个开放权重和三个封闭权重模型系列上进行评估。我们引入了三层分类法:(L1)直接单记录损坏、(L2)组合多记录损坏和(L3)上下文触发的休眠损坏。评估揭示了一个明显的防御边界:虽然基线写入时防御(如一致性检查)大幅抑制直接L1攻击,但无法可靠抑制L2和L3攻击。通过机制影响分解(MID),我们展示了写入时防御中的结构盲点,这些盲点允许看似良性的记录通过联合检索组合或触发条件激活后来变得有害。我们的发现主张从静态过滤转向自适应、上下文敏感的内存防御策略。

英文摘要

Persistent external memory enhances agent continuity but introduces persistent security vulnerabilities: adversarial content can be injected via standard interaction channels, retained across turns, and later distort downstream behavior. To address this challenge, we propose MemPoison, a comprehensive benchmark and analysis framework featuring 1227 hand-validated cases across four attack types, three injection channels, and three representative memory substrates, evaluated on seven open-weight and three closed-weight model families. We introduce a three-tier taxonomy: (L1) direct single-record corruption, (L2) compositional multi-record corruption and (L3) context-triggered dormant corruption. Our evaluations reveal a distinct defense frontier: while baseline write-time defenses, such as consistency checks, substantially suppress direct L1 attacks, they fail to reliably suppress L2 and L3 attacks. Through mechanistic influence decomposition (MID), we demonstrate structural blind spots in write-time defenses, which admit seemingly benign records that later become harmful through joint retrieval composition or trigger-conditioned activation. Our findings advocate for shifting from static filtering to adaptive, context-sensitive memory defense strategies.

URL PDF HTML 收藏
2607.00724 2026-07-17 cs.CL 版本更新

MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

MSQA:一个原生来源的多语言多文化SimpleQA基准

Xianru Chen, Yukai Huang, Mingxiang Chen, Xinping Lei, Fangbing Deng, Jin Chen, Ge Zhang, Wenhao Huang, Jiaheng Liu

机构 * M-A-P ByteDance Seed(字节跳动Seed) Beijing University of Posts and Telecommunications(北京邮电大学) Nanjing University(南京大学)

AI总结 提出MSQA基准,包含1064个原生问题覆盖11种语言和5个文化维度,评估18个LLM发现文化能力随预训练暴露程度下降,且推理时补救措施无效。

Comments Due to the company's data approval issue, we need to withdraw the article

详情
AI中文摘要

多语言流利性常常引发一个更强的假设:一个能说用户语言的模型也必须理解该语言编码的文化。我们称之为文化对齐的幻觉。为了直接检验这一假设,我们引入了MSQA,一个包含1064个原生来源问题的基准,涵盖11个语言组、五个文化维度和三个难度层级。与翻译基准不同,MSQA针对本地化知识,并减少了来自以英语为中心的跨语言迁移的捷径。评估18个LLM,我们发现显著的文化退化以及明显的局部性效应:文化能力更紧密地追踪预训练暴露程度,而非一般推理能力。我们进一步表明,常见的推理时补救措施并不能消除这种幻觉。模型对不熟悉的文化问题仍然过度自信,重复采样产生不稳定而非可靠的正确性,检索增强对长尾事实的帮助不均匀。这些发现表明,文化对齐不能仅从多语言能力推断,并且需要比推理时的校准、采样或检索更深入的干预。

英文摘要

Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language. We call this the Illusion of Cultural Alignment. To test this assumption directly, we introduce MSQA, a benchmark of 1,064 natively sourced questions across 11 language groups, five cultural dimensions, and three difficulty tiers. Unlike translated benchmarks, MSQA targets locally grounded knowledge and reduces shortcuts from English-centric cross-lingual transfer. Evaluating 18 LLMs, we find substantial cultural degradation and a pronounced Locality Effect: cultural competence tracks pre-training exposure more closely than general reasoning ability. We further show that common inference-time remedies do not dissolve the illusion. Models remain overconfident on unfamiliar cultural questions, repeated sampling yields unstable rather than reliable correctness, and retrieval augmentation helps unevenly on long-tail facts. These findings indicate that cultural alignment cannot be inferred from multilingual ability alone and requires deeper intervention than calibration, sampling, or retrieval at inference time

URL PDF HTML 收藏
2607.13059 2026-07-16 cs.RO 新提交

GPUSimBench: Towards Scalable and Reliable GPU-Accelerated Simulators in Embodied AI

GPUSimBench:迈向具身人工智能中可扩展且可靠的GPU加速模拟器

Huzhenyu Zhang, Shenghai Yuan, Wenrui Yan, Li Ma, Hengjie Li, Jingcheng Pang, Dmitry Yudin

机构 * Shanghai AI Laboratory(上海人工智能实验室) MIRAI(未来人工智能研究所) Nanyang Technological University(南洋理工大学) Nanjing University(南京大学)

AI总结 研究具身人工智能中GPU加速模拟器问题,通过GPUSimBench工具,建立物理基础评估、基准测试并行可扩展性,揭示并量化GPU批处理执行的不确定性,确定模拟器堆栈随机经验模式,强调无界扩展对可重复性的影响。

Comments Accepted by IROS 2026

详情
AI中文摘要

数据驱动的具身人工智能正迅速转变为通过大规模并行模拟来扩展训练的范式,其中GPU加速模拟器是基础数据基础设施。然而,随着计算吞吐量的扩展,并行效率、物理保真度和执行确定性之间的潜在权衡仍未得到充分研究,阻碍了可靠机器人学习的发展。本文通过引入GPUSimBench揭示了主流基于GPU的机器人模拟器(如Isaac Lab、Genesis)的隐藏局限性,该工具专注于可扩展性、物理一致性和计算确定性。首先,GPUSimBench通过一个受控的斜面任务建立物理基础评估,量化模拟动力学与其现实世界对应物之间的分布对齐。其次,我们通过测量跨扩展环境数量的吞吐量和内存占用情况来基准测试并行可扩展性。至关重要的是,除了标准性能指标外,我们还揭示并量化了GPU批处理执行引入的固有不确定性,其特征是即使在相同初始条件下,运行到运行和环境间也存在显著差异。最后,我们确定了当前模拟器堆栈中的四种随机经验模式,强调无界扩展在没有明确约束时会损害可重复性。

英文摘要

Data-driven embodied AI is rapidly transitioning into a paradigm that scales training through massively parallel simulation, where GPU-accelerated simulators serve as the foundational data infrastructure. However, as computational throughput scales, the underlying trade-offs between parallel efficiency, physical fidelity, and execution determinism remain largely unexamined, hindering the development of reliable robot learning. In this paper, we expose the hidden limits of mainstream GPU-based robotic simulators (e.g., Isaac Lab, Genesis) by introducing GPUSimBench, which focuses on scalability, physical consistency, and computational determinism. First, GPUSimBench establishes a physical grounding evaluation with a controlled inclined-plane task, quantifying the distributional alignment between simulated dynamics and their real-world counterparts. Second, we benchmark parallel scalability by measuring throughput and memory footprints across scaling environment counts. Crucially, beyond standard performance metrics, we unveil and quantify the inherent non-determinism introduced by GPU-batched execution, characterized by significant run-to-run and inter-environment variability even under identical initial conditions. Finally, we identify four empirical regimes of stochasticity within current simulator stacks, highlighting that unbounded scaling can compromise reproducibility without explicit constraints.

URL PDF HTML 收藏
2606.15587 2026-07-16 cs.RO 版本更新

Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments

完美演示造就差劲教师:从关键运动片段学习鲁棒对齐

Mingyu Liu, Zeju Li, Jiuhe Shu, Hanqing Wang, Yuhao Chao, Hao Chen, Chunhua Shen

机构 * Zhejiang University(浙江大学) Shanghai Innovation Institute(上海创新研究院) Hong Kong University of Science and Technology (GZ)(香港科技大学(广州)) Nanjing University(南京大学)

AI总结 针对精细操作中流畅演示因压缩关键对齐动作导致策略学习不足的问题,提出数据级重采样和表示级STAIR特征,利用稠密运动感知监督提升策略鲁棒性。

详情
AI中文摘要

专家演示被广泛认为是机器人模仿学习的黄金标准。然而,对于插入、堆叠和对齐等精细操作,我们发现了一个反直觉的失败模式:流畅的演示可能是差劲的教师。熟练的遥操作员将对齐和恢复的决定性时刻压缩到一个短暂的时间窗口内,导致策略被冗余的自由空间运动淹没,并在精度决定成功的关键区域缺乏监督。我们在两个层面解决这一瓶颈。在数据层面,靠近对齐时减速和对关键片段重采样都有帮助,但收益主要来自拓宽策略必须学习的恢复状态的覆盖范围,而非重新加权已有的帧。然而,这种数据层面的修复并未触及策略的逐帧视角:单个图像仍然直接映射到动作,控制修正的局部运动仍然隐式。因此,我们转向表示层面,引入STAIR(时空特征作为机器人学习接口),这是一种紧凑的动态特征,连接视觉-语言模型和动作专家,将每个轨迹中已记录的短视运动蒸馏为稠密的、运动感知的监督。仅使用流畅数据训练,STAIR恢复了大部分精心演示带来的增益(总体从50.0%提升至62.2%,接近精心演示的64.4%)。这些结果呼吁对机器人数据采取更具教学性的视角,优化机器的可学习性而非仅考虑人类效率。

英文摘要

Expert demonstrations are widely assumed to be the gold standard for robot imitation learning. Yet for fine-grained manipulation such as insertion, stacking, and alignment, we uncover a counterintuitive failure mode: fluent demonstrations can be poor teachers. A skilled teleoperator compresses the decisive moments of alignment and recovery into a brief temporal window, leaving the policy flooded with redundant free-space motion and starved of supervision exactly where precision determines success. We address this bottleneck at two levels. At the data level, slowing down near alignment and resampling critical segments both help, yet the gain comes mainly from broadening the coverage of recovery states the policy must learn, not from reweighting frames it already has. Such data-side fixes, however, leave the policy's per-frame view untouched: a single image still maps directly to an action, and the local motion that governs correction stays implicit. We therefore turn to the representation level and introduce STAIR (\textbf{S}patio-\textbf{T}emporal feature \textbf{A}s an \textbf{I}nterface for \textbf{R}obot learning), a compact dynamic feature that bridges the vision-language model and the action expert, distilling the short-horizon motion already recorded in each trajectory into dense, motion-aware supervision. Trained on fluent data alone, STAIR recovers most of the deliberate-demonstration gain ($50.0$ to $62.2\%$ overall, approaching the $64.4\%$ of deliberate demonstrations). These results call for a more pedagogical view of robot data, optimized for machine learnability rather than human efficiency alone.

URL PDF HTML 收藏
2603.16307 2026-07-16 cs.AI 版本更新

NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing

NeSy-Route:一种用于遥感约束路径规划的神经符号基准

Ming Yang, Zhi Zhou, Shi-Yu Tian, Kun-Yang Yu, Lan-Zhe Guo, Yu-Feng Li

机构 * State Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学) School of Artifical Intelligence, Nanjing University, China(人工智能学院,南京大学) School of Intelligence Science and Technology, Nanjing University, China(智能科学与技术学院,南京大学)

AI总结 NeSy-Route是一个大规模神经符号基准,用于评估遥感中的约束路径规划能力,通过自动化数据生成框架和三级评估协议,全面测试多模态大语言模型的感知、推理和规划能力。

Comments Accepted by ECCV2026

详情
AI中文摘要

遥感技术支撑着灾害救援和生态调查等关键应用,其中系统必须理解复杂场景和约束并做出可靠决策。当前遥感基准主要评估多模态大语言模型(MLLMs)的感知和推理能力,但未能评估规划能力,原因在于大规模定制和验证规划任务的难度或评估协议不准确。为解决这些限制,我们引入NeSy-Route,一个大规模神经符号基准,用于遥感中的约束路径规划。在此基准中,我们引入了一个自动化数据生成框架,整合高保真语义掩码与启发式搜索,生成具有可证明最优解的多样化路径规划任务。这使NeSy-Route能够全面评估10,821个路径规划样本,接近现有最大基准的十倍。此外,开发了一种三级分层神经符号评估协议,以实现准确评估并支持对感知、推理和规划的同时细粒度分析。我们对各种最先进的MLLMs进行了全面评估,表明现有MLLMs在感知和规划能力上存在显著缺陷。我们希望NeSy-Route能支持进一步研究和开发更强大的MLLMs用于遥感。

英文摘要

Remote sensing underpins crucial applications such as disaster relief and ecological field surveys, where systems must understand complex scenes and constraints and make reliable decisions. Current remote-sensing benchmarks mainly focus on evaluating perception and reasoning capabilities of multimodal large language models (MLLMs). They fail to assess planning capability, stemming either from the difficulty of curating and validating planning tasks at scale or from evaluation protocols that are inaccurate and inadequate. To address these limitations, we introduce NeSy-Route, a large-scale neuro-symbolic benchmark for constrained route planning in remote sensing. Within this benchmark, we introduce an automated data-generation framework that integrates high-fidelity semantic masks with heuristic search to produce diverse route-planning tasks with provably optimal solutions. This allows NeSy-Route to comprehensively evaluate planning across 10,821 route-planning samples, nearly 10 times larger than the largest prior benchmark. Furthermore, a three-level hierarchical neuro-symbolic evaluation protocol is developed to enable accurate assessment and support fine-grained analysis on perception, reasoning, and planning simultaneously. Our comprehensive evaluation of various state-of-the-art MLLMs demonstrates that existing MLLMs show significant deficiencies in perception and planning capabilities. We hope NeSy-Route can support further research and development of more powerful MLLMs for remote sensing. The dataset and code are available at https://mingyang1010.github.io/NeSy-Route/.

URL PDF HTML 收藏
2603.12712 2026-07-16 cs.SE cs.LG 版本更新

Design-Specification Tiling for ICL-based CAD Code Generation

基于ICL的CAD代码生成的Design-Specification Tiling设计

Yali Du, San-Zhuo Xi, Hui Sun, Ming Li

机构 * National Key Laboratory for Novel Software Technology, School of Artificial Intelligence, Nanjing University(新型软件技术国家实验室,人工智能学院,南京大学)

AI总结 本文提出DST方法,通过量化知识充分性提升CAD代码生成质量,优于现有ICL示例选择策略。

详情
AI中文摘要

Large language models (LLMs) have demonstrated remarkable capabilities in code generation, yet they underperform on domain-specific tasks such as Computer-Aided Design (CAD) code generation due to scarce training data. In-Context Learning (ICL) offers a 免训练 alternative through task-specific exemplars. However, existing selection strategies prioritize similarity or point-wise diversity, often producing redundant selections that fail to satisfy the compositional requirements of complex CAD design specifications. In this work, we propose knowledge sufficiency as a principled objective for exemplar selection that aims to maximally satisfy all requirements within design specifications. To realize this objective, we introduce Design-Specification Tiling (DST), which quantifies knowledge sufficiency through a surrogate tiling ratio by extracting multi-granular design components and measuring the proportion of query components covered by selected exemplars. We demonstrate that maximizing this objective constitutes submodular maximization and provide a polynomial-time greedy algorithm with a (1-1/e)-approximation guarantee. Extensive experiments demonstrate that DST substantially improves CAD code generation quality, consistently outperforming existing exemplar selection strategies in ICL.

英文摘要

Large language models~(LLMs) have demonstrated remarkable capabilities in code generation, yet their performance remains limited on domain-specific tasks such as Computer-Aided Design~(CAD) code generation, largely due to the scarcity of high-quality training data. In-Context Learning~(ICL) provides a training-free alternative by prompting LLMs with task-specific exemplars, but its effectiveness critically depends on how exemplars are selected. Existing selection strategies mainly rely on similarity or point-wise diversity, often overlooking the compositional nature of CAD design specifications, where a query may involve multiple functional requirements, geometric constraints, and design primitives. As a result, selected exemplars can be individually relevant but collectively redundant, providing insufficient coverage for complex design requirements. In this work, we propose \emph{knowledge sufficiency} as a principled objective for exemplar selection, aiming to select a compact set of exemplars that maximally satisfies the requirements contained in a target design specification. To instantiate this objective, we introduce \emph{Design-Specification Tiling~(DST)}, which estimates knowledge sufficiency through a surrogate tiling ratio by decomposing design specifications into multi-granular components and measuring the proportion of query components covered by selected exemplars. We further show that optimizing this objective can be formulated as a submodular maximization problem, and develop a polynomial-time greedy algorithm tailored to this setting with a $(1-1/e)$-approximation guarantee. Extensive experiments across multiple LLMs demonstrate that DST substantially improves CAD code generation quality and consistently outperforms existing ICL exemplar selection strategies, highlighting the importance of requirement-level knowledge coverage for domain-specific code generation.

URL PDF HTML 收藏
2601.12349 2026-07-16 cs.CR cs.AI cs.SE 版本更新

Mind the Gap: Action Rebinding Attacks against Android GUI Agents

零权限操纵:我们能否信任由大型多模态模型驱动的GUI代理?

Yi Qian, Kunwei Qian, Xingbang He, Ligeng Chen, Jikang Zhang, Tiantai Zhang, Haiyang Wei, Linzhang Wang, Hao Wu, Bing Mao

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) Hornor Device Co., Ltd(Hornor设备有限公司) Institute of Dataspace, Hefei Comprehensive National Science Center(数据空间研究所,合肥综合性国家科学中心)

AI总结 研究揭示了由大型多模态模型驱动的GUI代理在Android中存在视觉原子性假设的漏洞,通过零权限攻击实现对代理执行的重新绑定,并利用意图对齐策略绕过验证门,展示了代理-操作系统集成中的安全缺陷。

详情
AI中文摘要

大型多模态模型驱动的GUI代理正在移动平台上作为高权限操作者出现,负责感知屏幕内容并注入输入。然而,其设计基于隐含的视觉原子性假设:即UI状态在观察和行动之间保持不变。我们证明在Android中这一假设从根本上不成立,从而创建了一个关键的攻击面。我们提出了Action Rebinding攻击,一种允许一个看似无害的应用程序利用零危险权限重新绑定代理执行的攻击方式。通过利用代理推理管道中固有的观察到行动的间隙,攻击者触发前台转换,将代理的计划行动重新定向至目标应用程序。我们利用代理的任务恢复逻辑和Android的UI状态保存来编排可编程的多步骤攻击链。此外,我们引入了意图对齐策略(IAS),该策略操纵代理的推理过程以合理化UI状态,使其能够绕过验证门(例如确认对话框)否则会被拒绝。我们对六个广泛使用的Android GUI代理在15个任务上的Action Rebinding攻击进行了评估。我们的结果表明,对于原子动作重新绑定,成功率为100%,并且能够可靠地编排多步骤攻击链。借助IAS,绕过验证门的成功率提高(从0%到最高100%)。值得注意的是,攻击应用程序不需要敏感权限且不包含任何特权API调用,实现了跨恶意软件扫描器(例如VirusTotal)的0%检测率。我们的发现揭示了当前代理-操作系统集成中的基本架构缺陷,并为未来代理系统的安全设计提供了关键见解。如需访问实验日志和演示视频,请联系yi_qian@smail.nju.edu.cn。

英文摘要

Large multimodal model powered GUI agents are emerging as high-privilege operators on mobile platforms, entrusted to perceive screen content and inject inputs across application boundaries. While these agents aim to automate complex tasks, we demonstrate that their design introduces a fundamental conflict with Android's strict application sandboxing. We present a novel cross-application Action Rebinding attack, which allows a malicious application with zero dangerous permissions to hijack the agent's execution and perform privileged operations on behalf of the attacker. Our attack exploits the inevitable observation-action gap inherent in the agent's reasoning pipeline. A malicious app can render a benign ``contextual carrier'' to elicit a planned action, and then swap the foreground to a sensitive target application during the reasoning latency. The agent, unaware of the transition, unwittingly executes the action in the privileged context. We further advance this attack by weaponizing the agent's own task-recovery logic to create programmable, multi-step exploit loops , and introducing an Intent Alignment Strategy (IAS) that manipulates the agent's reasoning to rationalize the hijacked state. We evaluate our attack on six widely-used Android GUI agents. Our results demonstrate a 100% success rate for atomic action hijacking and the ability to orchestrate high-impact exploits, including unauthorized file deletion, SMS transmission, and app uninstallation, without the attacker holding any corresponding permissions. Furthermore, since the malicious application separates intent from capability and contains no privileged API calls, it achieves a 0% detection rate across commercial malware scanners (e.g., VirusTotal), highlighting a critical blind spot in current mobile security analysis. To access experimental logs and demonstration videos, please contact yi_qian@smail.nju.edu.cn.

URL PDF HTML 收藏
2505.09433 2026-07-16 cs.CV eess.IV 版本更新

Efficient LiDAR Reflectance Compression via Scanning Serialization

通过扫描序列化实现高效的激光雷达反射率压缩

Jiahao Zhu, Kang You, Dandan Ding, Zhan Ma

机构 * School of Information Science and Technology, Hangzhou Normal University, Hangzhou, China(信息科学与技术学院,杭州师范大学) School of Electronic Science and Engineering, Nanjing University, Nanjing, China(电子科学与工程学院,南京大学)

AI总结 研究针对激光雷达反射率在神经压缩方法中未充分探索的问题,提出基于序列化的SerLiC框架。通过扫描序列化转换点云、标记点及采用双并行化方案建模,实验表明其压缩效果优,轻量级版本性能好,对实际应用有吸引力。

详情
AI中文摘要

激光雷达点云中的反射率属性为下游任务提供重要信息,但在神经压缩方法中未得到充分探索。为解决此问题,我们引入了SerLiC,这是一种基于序列化的神经压缩框架,以充分利用激光雷达反射率的内在特性。SerLiC首先通过扫描顺序序列化将3D激光雷达点云转换为1D序列,为反射率分析提供以设备为中心的视角。然后将每个点标记为包含其传感器扫描索引、径向距离和先前反射率的上下文表示,以有效探索依赖性。为了进行高效的顺序建模,Mamba采用了双并行化方案,能够同时捕获自回归依赖性并进行快速处理。大量实验表明,SerLiC相对于原始反射率数据实现了超过2倍的体积减少,在仅使用其2%参数的情况下,压缩比特减少比现有方法高出22%。此外,轻量级版本的SerLiC仅用111K参数就能实现>10帧/秒,对实际应用具有吸引力。

英文摘要

Reflectance attributes in LiDAR point clouds provide essential information for downstream tasks but remain underexplored in neural compression methods. To address this, we introduce SerLiC, a serialization-based neural compression framework to fully exploit the intrinsic characteristics of LiDAR reflectance. SerLiC first transforms 3D LiDAR point clouds into 1D sequences via scan-order serialization, offering a device-centric perspective for reflectance analysis. Each point is then tokenized into a contextual representation comprising its sensor scanning index, radial distance, and prior reflectance, for effective dependencies exploration. For efficient sequential modeling, Mamba is incorporated with a dual parallelization scheme, enabling simultaneous autoregressive dependency capture and fast processing. Extensive experiments demonstrate that SerLiC attains over 2x volume reduction against the original reflectance data, outperforming the state-of-the-art method by up to 22% reduction of compressed bits while using only 2% of its parameters. Moreover, a lightweight version of SerLiC achieves > 10 fps (frames per second) with just 111K parameters, which is attractive for real-world applications.

URL PDF HTML 收藏
2607.12820 2026-07-15 cs.CV 新提交

AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning

AVSCap:为全模态视频字幕编排视听协同

Yanghai Wang, Jiahao Wang, Jiafu Tang, Yuanxing Zhang, Zhe Cao, Hanyan Bian, Zijie Zhang, Weiliang Luo, Zhiyu Pan, Zixuan Dong, Jiaheng Liu, Zhaoxiang Zhang

机构 * Nanjing University(南京大学) Kuaishou Technology(快手科技) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

AI总结 研究全模态视频字幕编排视听协同问题,提出AVSCap框架,构建训练语料库,采用两阶段策略训练字幕生成器,引入新基准,实验表明该模型在非语音音频覆盖率和跨模态绑定方面表现出色,提升了整体性能。

详情
AI中文摘要

全模态视频字幕不仅仅是将视觉字幕与音频转录相结合:一个有用的字幕必须描述视觉动作、语音、音乐和音效如何共同演变。现有的大型多模态模型在这个关系步骤上常常失败,将音频和视觉流视为松散耦合的观察结果,依赖自动语音识别,并且未充分说明非语音声音及其与视觉事件的联系。我们提出了AVSCap,一个以明确的跨模态事件绑定为中心的视听字幕框架。首先,我们构建了AVSCap-130K,这是一个通过解耦然后融合的管道生成的三模态训练语料库,在组合有基础的全模态字幕之前锚定视觉和声学证据。其次,我们训练了AVSCap-7B,一个具有两阶段策略的7B字幕生成器:监督微调建立基线能力,而样本高效的强化学习使用混合奖励来优化声学完整性和视听协同。我们的缩放分析表明,强化学习比增加监督微调数据带来更大的收益。第三,我们引入了AVSCapBench基准,该基准将字幕分解为视觉、音频和协同事件,并使用细粒度事件召回率对其进行评估。在AVSCapBench和外部基准上的实验表明,AVSCap-7B提高了非语音音频覆盖率和跨模态绑定,在评估的开源模型中提供了最佳的整体性能。

英文摘要

Omni-modal video captioning is not merely combining visual captioning with audio transcription: a useful caption must describe how visual actions, speech, music, and sound effects co-evolve. Existing large multimodal models often fail at this relational step, treating audio and visual streams as loosely coupled observations, relying on automatic speech recognition, and under-specifying non-speech sounds and their links to visual events. We present AVSCap, a framework for audio-visual captioning centered on explicit cross-modal event binding. First, we construct AVSCap-130K, a tri-modal training corpus generated by a decoupled-then-fused pipeline that anchors visual and acoustic evidence before composing grounded omni-modal captions. Second, we train AVSCap-7B, a 7B captioner with a two-stage strategy: supervised fine-tuning establishes baseline capabilities, while sample-efficient reinforcement learning uses hybrid rewards to optimize acoustic completeness and audio-visual synergy. Our scaling analysis shows that reinforcement learning brings larger gains than increasing SFT data. Third, we introduce AVSCapBench, a benchmark that decomposes captions into visual, audio, and synergy events and evaluates them with fine-grained event recall. Experiments on AVSCapBench and external benchmarks show that AVSCap-7B improves non-speech audio coverage and cross-modal binding, delivering the best overall performance among evaluated open-source models.

URL PDF HTML 收藏
2607.10183 2026-07-15 cs.DC cs.AI cs.LG 版本更新

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

用于消费级设备上混合CPU - GPU大语言模型推理的自动张量调度

Yangyijian Liu, Hongyi Ye, Mingyang Li, Wu-jun Li

机构 * School of Computer Science, Nanjing University(南京大学计算机科学学院)

AI总结 研究在消费级设备运行大语言模型时因模型权重超GPU内存需卸载推理的问题,提出ATSInfer系统,通过结合静态张量放置与负载感知动态传输及异步协调,在代表性平台评估,相比现有系统显著提升吞吐量、利用率等,改善本地大语言模型部署体验。

详情
AI中文摘要

在笔记本电脑和台式机等消费级设备上运行大语言模型具有挑战性,因为模型权重常超GPU内存容量,需借助CPU内存进行卸载推理。现有卸载系统依赖粗略调度,忽略张量异质性且难适应硬件负载变化。本文提出ATSInfer,一种在张量粒度上执行卸载的混合CPU - GPU推理系统。它结合静态张量放置与负载感知动态传输,引入异步CPU - GPU协调,跨异构后端高效调度硬件存储、数据移动和计算。在代表性消费平台上用密集和MoE模型评估,相比现有系统,ATSInfer提高预填充吞吐量达1.94倍,解码吞吐量达3.29倍,还提升了GPU利用率并更有效利用PCIe带宽,显著改善了个人消费设备上本地大语言模型部署的用户体验。

英文摘要

Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference necessary to extend effective model capacity with CPU memory. Existing offloading systems, however, typically rely on coarse layer-level or expert-level scheduling, which overlooks substantial heterogeneity among tensors within the same layer and adapts poorly to changing hardware load conditions on such devices. This paper presents ATSInfer, a hybrid CPU-GPU inference system for consumer devices that performs offloading at tensor granularity. ATSInfer combines static tensor placement with load-aware dynamic transfer, and introduces asynchronous CPU-GPU coordination to efficiently schedule hardware storage, data movement, and computation across heterogeneous backends. We implement ATSInfer and evaluate it on representative consumer platforms using both dense and MoE models. Compared with existing systems, ATSInfer improves prefill throughput by up to 1.94$\times$ and decode throughput by up to 3.29$\times$, while also increasing GPU utilization and making more effective use of PCIe bandwidth. These results show that ATSInfer can substantially improve the user experience of local LLM deployment on personal consumer devices.

URL PDF HTML 收藏
2602.05513 2026-07-15 cs.RO cs.AI 版本更新

DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter

DECO:解耦多模态扩散变压器用于配备插件触觉适配器的双臂灵巧操作

Xukun Li, Yu Sun, Lei Zhang, Bosheng Huang, Yibo Peng, Yuan Meng, Haojun Jiang, Shaoxuan Xie, Guocai Yao, Alois Knoll, Zhenshan Bing, Xinlong Wang, Zhenguo Sun

机构 * Beijing Academy of Artificial Intelligence, Beijing, China(北京人工智能研究院) School of Computation, Information and Technology, Technical University of Munich, Garching, Germany(慕尼黑技术大学计算与信息学院) Department of Shenyang Institute of Computing Technology, University of Chinese Academy of Sciences, Beijing, China(中国科学院沈阳计算技术研究所部门) State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China(新型软件技术国家重点实验室) Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系)

AI总结 DECO通过解耦多模态输入和触觉适配器,实现了双臂灵巧操作的高效整合与高成功率

Comments 17 pages, 8 figures. Project Page: https://baai-humanoid.github.io/DECO-webpage/

详情
AI中文摘要

双臂灵巧操作依赖于整合多模态输入以执行复杂的现实任务。为了解决有效结合这些模态的挑战,我们提出了DECO,一种解耦的多模态扩散变压器,通过专门的条件路径解耦视觉、本体感觉和触觉信号,实现多模态输入的结构化和可控整合,并配备轻量级适配器以高效注入额外信号。此外,我们发布了DECO-50数据集,用于配备触觉传感的双臂灵巧操作,包含50小时的数据和超过500万帧,通过远程操作在真实的双臂机器人上收集。我们训练DECO在DECO-50上,并进行了超过2000次机器人运行的广泛现实评估。实验结果表明,DECO在所有任务上均取得最佳性能,平均成功率为72.25%,比基线提高了21%。此外,触觉适配器在所有任务上带来了额外的10.25%平均成功率,并在复杂接触丰富的任务上获得20%的提升,同时仅调整了模型参数的不到10%。

英文摘要

Bimanual dexterous manipulation relies on integrating multimodal inputs to perform complex real-world tasks. To address the challenges of effectively combining these modalities, we propose DECO, a decoupled multimodal diffusion transformer that disentangles vision, proprioception, and tactile signals through specialized conditioning pathways, enabling structured and controllable integration of multimodal inputs, with a lightweight adapter for parameter-efficient injection of additional signals. Alongside DECO, we release DECO-50 dataset for bimanual dexterous manipulation with tactile sensing, consisting of 50 hours of data and over 5M frames, collected via teleoperation on real dual-arm robots. We train DECO on DECO-50 and conduct extensive real-world evaluation with over 2,000 robot rollouts. Experimental results show that DECO achieves the best performance across all tasks, with a 72.25% average success rate and a 21% improvement over the baseline. Moreover, the tactile adapter brings an additional 10.25% average success rate across all tasks and a 20% gain on complex contact-rich tasks while tuning less than 10% of the model parameters.

URL PDF HTML 收藏
2607.11124 2026-07-14 cs.SD cs.AI 新提交

BeatEdit: Symbolic Music Generation as Explicit Editing

BeatEdit:作为显式编辑的符号音乐生成

Haoyu Gu, Lekai Qian, Haowu Zhou, Qi Liu, Shuai Wang

机构 * School of Future Technology South China University of Technology Guangzhou China(未来技术学院 华南理工大学 广州 中国) South China University of Technology(华南理工大学) Nanjing University(南京大学)

AI总结 研究针对符号音乐生成中缺乏选择性修改支持的问题,提出BeatEdit框架,基于BEAT编码,含三种互补机制,共享单一编码和预训练主干,在多项任务中精度和质量更高且高效,揭示编码设计对编辑有效性有重要影响。

详情
AI中文摘要

音乐创作本质上是一个修订过程。然而,符号音乐生成仍主要由从头生成完整序列的范式主导,对选择性修改的支持有限。基于编辑的方法在文本转换任务中已证明有效,但在符号音乐方面基本未被探索。我们将这种缺失追溯到表示层面:传统的基于事件的音乐编码缺乏显式音乐编辑所需的结构属性。相比之下,BEAT编码是一种最初为自回归生成设计的基于节拍网格的表示,具有适合编辑的结构属性。我们提出了BeatEdit,这是第一个基于显式编辑操作的符号音乐生成框架,将生成重新定义为通过编辑草稿而不是从头合成来产生新内容。BeatEdit包括沿编辑密度增加轴的三种互补机制:用于纠错的逐令牌序列标记、用于伴奏编辑的迭代细化以及用于片段完成 的标记然后填充。所有这些机制共享单一编码和预训练主干,在所有三项任务中比自回归和扩散方法实现更高的精度和感知质量,同时保持高效,单通道推理在100毫秒内完成。交叉编码评估进一步表明,编码设计对编辑有效性有重大影响,存在显著的编码方法交互效应。代码可在这个https URL获取。

英文摘要

Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete sequences from scratch, with limited support for selective modification. Edit-based methods have proven effective for text transformation tasks, but remain largely unexplored for symbolic music. We trace this absence to the representational level: conventional event-based music encodings lack the structural properties required by explicit music editing. In contrast, the BEAT encoding, a beat-grid-anchored representation originally designed for autoregressive generation, possesses structural properties amenable to editing. We propose BeatEdit, the first framework for symbolic music generation based on explicit edit operations, recasting generation as producing new content by editing a draft rather than synthesizing from scratch. BeatEdit comprises three complementary mechanisms along an axis of increasing edit density: per-token sequence tagging for error correction, iterative refinement for accompaniment editing, and tag-then-fill for segment completion. All these mechanisms share a single encoding and pre-trained backbone, achieving higher precision and perceptual quality than autoregressive and diffusion methods across all three tasks, while remaining efficient, with single-pass inference completing in under 100 ms. Cross-encoding evaluation further reveals that encoding design substantially influences editing effectiveness, with notable encoding-method interaction effects. Code is available at https://github.com/Haoyu-Gu/BeatEdit-code

URL PDF HTML 收藏
2607.11008 2026-07-14 cs.CV cs.AI 新提交

SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception

SynCLIP:用于鲁棒开放词汇密集感知的同义词一致语言-图像预训练

Mingjie Xie, Guangjun He, Dongli Xu, Youtian Lin, Hongjue Li, Pengming Feng, Jian Guan, Yue Deng

机构 * Beihang University(北京航空航天大学) State Key Laboratory of Space Information System and Integrated Application(空间信息系统与集成应用国家重点实验室) Nanjing University(南京大学) Harbin Engineering University(哈尔滨工程大学) Beijing Zhongguancun Academy(北京中关村科学城)

AI总结 研究针对开放词汇密集感知中同义词引起的定位不一致问题,提出SynCLIP框架,通过SSA和SAR模块增强注意力一致性与定位精度,构建SEViC支持预训练,实验证明该方法显著提升定位一致性并达领先性能。

Comments Accepted by CVPR 2026

详情
AI中文摘要

开放词汇密集感知(OVDP)旨在通过利用文本知识来定位训练期间未见过的物体。尽管基于CLIP的方法最近取得了显著进展,但存在一个关键限制:同义词引起的定位不一致,即语义等效的表达式会产生不同的空间注意力模式。这种不一致削弱了现有方法在实际OVDP应用中的鲁棒性和性能。为解决此问题,我们提出了SynCLIP,一个同义词一致语言-图像预训练框架,用于增强OVDP的同义词鲁棒定位。SynCLIP引入了语义一致空间注意力对齐(SSA)模块,通过最小化原始和同义词表达式的注意力图之间的差异来增强空间注意力一致性。此外,空间注意力细化(SAR)模块在对齐图中选择性地加强最语义相关的空间区域,以实现更精确和稳定的定位。为支持同义词一致预训练,我们还构建了一个同义词丰富视觉语料库(SEViC),用多个同义词和文本定义扩充每个类别。在多个基准上的广泛实验表明,SynCLIP在不同语言变体下显著提高了定位一致性,并在基于CLIP的OVDP方法中取得了领先性能。

英文摘要

Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remarkable progress of recent CLIP-based approaches, we identify a critical limitation: synonym-induced grounding inconsistency, where semantically equivalent expressions yield disparate spatial attention patterns. This inconsistency undermines the robustness and performance of existing methods in real-world OVDP applications. To address this issue, we propose SynCLIP, a Synonym-Coherent Language-Image Pretraining framework that enhances synonym-robust grounding for OVDP. SynCLIP introduces a Semantic-consistent Spatial Attention alignment (SSA) module to enhance spatial attention consistency by minimizing discrepancies between attention maps of original and synonymous expressions. Furthermore, a Spatial Attention Refinement (SAR) module selectively strengthens the most semantically relevant spatial regions within aligned maps for more precise and stable grounding. To support synonym-coherent pretraining, we also construct a Synonym-Enriched Visual Corpus (SEViC), which augments each category with multiple synonyms and textual definitions. Extensive experiments on multiple benchmarks demonstrate that SynCLIP substantially improves grounding consistency under diverse linguistic variants and achieves state-of-the-art performance among CLIP-based OVDP methods. Code is available at https://github.com/Justlovesmile/SynCLIP.

URL PDF HTML 收藏
2607.10120 2026-07-14 cs.CV 新提交

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding

WeaveEarth:用于免训练超高分辨率遥感理解的结构化证据构建与推理

Xianzhi Ma, Shujun Wang, Xiaohan Li, Hao Liu, Changhua Pei, Jianhui li

机构 * School of Frontier Sciences, Nanjing University(南京大学前沿科学学院) Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心)

AI总结 针对UHR遥感图像理解,提出免训练框架WeaveEarth,通过全局感知证据构建选最小支持证据集,经结构化证据推理增强VLM全局-局部联合推理能力,在多基准测试中表现优于现有方法。

详情
AI中文摘要

超高分辨率(UHR)遥感图像理解要求视觉语言模型(VLM)在有限计算预算下捕捉全局场景布局和稀疏但关键的局部细节。现有方法主要有被动感知和主动感知两种范式,但都存在不足。本文提出WeaveEarth,一个免训练框架,将UHR理解重新表述为全局上下文约束下的结构化证据构建与推理问题。具体包括全局感知证据构建以选择最小支持证据集,以及结构化证据推理将局部证据等编织成统一推理接口,增强VLM全局-局部联合推理能力。实验表明WeaveEarth在多个基准测试中优于现有方法。

英文摘要

Ultra-High-Resolution (UHR) remote sensing image understanding requires Vision-Language Models (VLMs) to capture both the global scene layout and sparse yet task-critical local details under limited computational budgets. Existing methods mainly follow two paradigms. One is passive perception, which relies on resolution expansion or token compression and may therefore discard fine-grained details. The other is active perception, which depends on multi-round zooming and search, but suffers from high latency, contextual fragmentation, and error accumulation. We argue that a more effective path toward UHR understanding lies not in accessing more, but in organizing better. To this end, we propose WeaveEarth, a training-free framework that reformulates UHR understanding as a problem of structured evidence construction and reasoning under global context constraints. Specifically, WeaveEarth first employs Global-Aware Evidence Construction to select a compact, low-redundancy, and spatially complementary Minimal Support Evidence Set. It then introduces Structured Evidence Reasoning, which weaves local evidence, spatial metadata, and relative topology into a unified reasoning interface, thereby enhancing the VLM's ability to perform global-local joint reasoning. Extensive experiments show that WeaveEarth consistently outperforms strong baselines and existing UHR methods across multiple UHR remote sensing benchmarks and multiple frozen VLM backbones. Code is available at https://github.com/XianZhi-Ma/WeaveEarth.

URL PDF HTML 收藏
2607.10664 2026-07-14 stat.ML cond-mat.mtrl-sci cs.LG physics.chem-ph 新提交

Edge Cluster Expansion with Radial Rotary Attention for Interatomic Potentials

用于原子间势的具有径向旋转注意力的边簇扩展

Zemin Xu, Wenbo Xie, P. Hu

机构 * School of Physical Science and Technology, ShanghaiTech University, Shanghai, China(物理科学与技术学院,上海科技大学,上海,中国) School of Chemistry, Nanjing University, Nanjing, China(化学学院,南京大学,南京,中国)

AI总结 研究针对机器学习原子间势的SO(2)理论局限性,提出维格纳D矩阵构造及两个新相互作用构建块,包括边复积基和径向旋转复注意力,改进原子簇扩展模块,训练的模型在Matbench Discovery上达最优性能。

详情
AI中文摘要

在本文中,我们对用于机器学习原子间势(MLIPs)的SO(2)理论进行了系统研究,并确定了传统SO(2)线性架构相对于SO(3)克莱布什 - 戈尔丹张量积(CGTP)的局限性。基于这些见解,我们提出了维格纳D矩阵的直接笛卡尔构造和递归克莱布什 - 戈尔丹构造,并引入了两个新颖的相互作用构建块。首先,基于广义不对称收缩提出了边复积基,这是一种多体展开的新公式,通过复值等变乘法直接在边上构建高阶相互作用。其次,引入了径向旋转复注意力(RRA),它提高了外推性能并超越了现有的注意力向量公式。我们还对原子簇扩展模块进行了一些改进。在此基础上,我们在OMat24、sAlex和MPTrj上训练模型,并推出了TECE - OAM - RRA - 1.0,在Matbench Discovery上实现了最优性能。

英文摘要

In this paper, we provide a systematic investigation of SO(2) theory to machine learning interatomic potentials (MLIPs) and identify the limitations of conventional SO(2) Linear architectures relative to SO(3) Clebsch-Gordan Tensor Products (CGTP). Building on these insights, we propose direct Cartesian construction and recursive Clebsch-Gordan construction of Wigner D-matrices and introduce two novel interaction building blocks. First, we propose the Edge Complex Product Basis based on Generalized Asymmetric Contraction, a new formulation for many-body expansion that directly constructs higher-order interactions on edges through complex-valued equivariant multiplications. Second, we introduce Radial Rotary Complex Attention(RRA), which enhances extrapolation performance and surpasses existing attention vector formulations. We also introduce several improvements to the Atomic Cluster Expansion module. Building on these advances, we train our models on OMat24, sAlex, and MPTrj, and introduce TECE-OAM-RRA-1.0, which achieve state-of-the-art (SOTA) performance on the Matbench Discovery.

URL PDF HTML 收藏