arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Xi'an Jiaotong University(西安交通大学)

至 收录 769
2607.17504 2026-07-21 cs.CV cs.AI 新提交

DecoyFace: Beyond Obfuscation via Controllable and Imperceptible Identity Misdirection for Privacy-Preserving Face Recognition

DecoyFace:通过可控且不可察觉的身份误导实现超越混淆的隐私保护人脸识别

Zhihan Ren, Lijun He, Xinyao Wang, Xinzhu Fu, Fan Li

机构 * School of Information and Communications Engineering, Xi’an Jiaotong University(西安交通大学信息与通信工程学院) Shaanxi Key Laboratory of Deep Space Exploration Intelligent Information Technology(陕西省深空探测智能信息技术重点实验室)

AI总结 研究针对分割式人脸识别中隐私问题,提出DecoyFace框架,通过分解中间表示为子空间,客户端注入诱饵线索,服务器端规范化处理,在保持识别准确率的同时大幅降低身份泄露,实现隐私保护的人脸识别。

详情
AI中文摘要

分割式人脸识别减少了客户端计算,但会将中间特征暴露于特征反转攻击以及诚实但好奇(HBC)服务器的未经授权分析。现有隐私保护人脸识别方法主要旨在抵御未经授权的重建,通常生成的特征反转后结果明显退化,可能揭示保护的存在并引发自适应攻击。为解决此问题,我们提出DecoyFace,一个面向不可察觉诱饵的框架,在保留识别效用的同时,将未经授权的重建导向一个看似合理但错误的身份。关键思想是将中间表示分解为对重建敏感的子空间及其互补子空间。客户端将诱饵身份线索注入对重建敏感的子空间,而来自真实样本的有限识别相关证据保留在互补子空间中。在服务器端,一个授权的规范化模块抑制诱饵主导的组件并恢复一个有利于识别的表示。此设计解决了攻击者从截获特征进行的反转以及HBC服务器从规范化表示进行的重建问题。实验表明,DecoyFace在保持有竞争力的识别准确率的同时,在U-Net攻击下将身份泄露大幅降低至2.93%,在Flow-Matching攻击下降低至0.74%,同时产生视觉上合理且不可察觉的重建,在LFW数据集上的面部有效性超过99.78%。

英文摘要

Split face recognition reduces client-side computation but exposes intermediate features to feature inversion attacks and unauthorized analysis by honest-but-curious (HBC) servers. Existing privacy-preserving face recognition methods mainly aim to resist unauthorized reconstruction, typically producing features whose inversion yields visibly degraded results, which may reveal the existence of protection and motivate adaptive attacks. To address this issue, we propose DecoyFace, an imperceptible decoy-oriented framework that steers unauthorized reconstruction toward a plausible but incorrect identity while preserving recognition utility. The key idea is to decompose the intermediate representation into a reconstruction-sensitive subspace and its complementary subspace. The client injects decoy identity cues into the reconstruction-sensitive subspace, while limited recognition-relevant evidence from the true sample is retained in the complementary subspace. On the server side, an authorized canonicalization module suppresses decoy-dominant components and recovers a recognition-friendly representation. This design addresses both attacker-side inversion from intercepted features and HBC server-side reconstruction from canonicalized representations. Experiments show that DecoyFace preserves competitive recognition accuracy while substantially reducing identity leakage to 2.93% under U-Net attacks and 0.74% under Flow-Matching attacks while yielding visually plausible and imperceptible reconstructions, with over 99.78% face validity on LFW dataset.

URL PDF HTML 收藏
2607.16621 2026-07-21 cs.CL 新提交

From Memory to Skills: Evidence-Grounded Co-Evolution Governance for Long-Horizon LLM Agents

从记忆到技能:基于证据的长期大语言模型智能体协同进化治理

Bo Tang, Yang Zhang, Guomian Zhuang, Wenqiang Wei, Gaoyang Zheng, Lindong Xie, Yanchao Tan, Feiyu Xiong, Qingyu Yang, Edward Chung, Zhiyu li

机构 * University of Science and Technology of China(中国科学技术大学) Hong Kong Polytechnic University(香港理工大学) Fuzhou University(福州大学) Xi’an Jiaotong University(西安交通大学)

AI总结 针对长期大语言模型智能体记忆系统问题,提出无需训练的MSCE框架,将经验组织为多种形式,把证据支持的策略转化为可调用技能,引入反射加权值回填,实验证明其性能优于基线,有跨域转移性和终身进化能力。

Comments Submitted into EMNLP'2026

详情
AI中文摘要

现有的长期大语言模型智能体记忆系统通常将先前痕迹作为被动上下文检索,而非转化为可执行能力。本文提出MSCE,一个无需训练的记忆-技能协同进化框架,将智能体经验组织为有根据的步骤痕迹、可复用的程序策略和声明性环境认知。MSCE将具有正估计增益的证据支持的二级策略结晶为可调用技能,保留证据链接、适用边界等。还引入反射加权值回填,通过密集局部自反射传播稀疏终端反馈,以产生用于治理记忆和技能进化的证据校准痕迹值。在EvoAgentBench和LoCoMo上的实验表明,MSCE显著优于现有技术的技能增强和记忆驱动智能体基线,具有强大的跨域可转移性和终身进化能力。

英文摘要

Existing memory systems for long-horizon LLM agents often retrieve prior traces as passive context rather than converting them into executable capabilities. In this paper, we propose MSCE, a training-free Memory--Skill Co-Evolution framework that organizes agent experience into grounded step traces, reusable procedural policies, and declarative environmental cognition. MSCE crystallizes evidence-backed L2 policies with positive estimated gain into callable skills that retain evidence links, applicability boundaries, decision guidance, verification rules, and reliability estimates. It further introduces reflection-weighted value backfilling, which propagates sparse terminal feedback through dense local self-reflections to produce evidence-calibrated trace values for governing memory and skill evolution. Experiments on EvoAgentBench and LoCoMo demonstrate that MSCE significantly outperforms state-of-the-art skill-augmented and memory-driven agent baselines, exhibiting strong cross-domain transferability and lifelong-evolution capabilities.

URL PDF HTML 收藏
2607.16599 2026-07-21 cs.SD cs.MM eess.AS 新提交

Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves

一个分数够吗?用时间分数曲线评估歌曲的演唱质量

Yishan Lv, Jing Luo, Xinyu Yang, Zhizheng Wu

机构 * Xi’an Jiaotong University(西安交通大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

AI总结 研究针对全长歌曲演唱质量评估难题,提出SongSQA两阶段框架。第一阶段用伪标签训练段分数预测器,第二阶段聚合器整合特征与分数生成嵌入并捕捉联系,有效提升评估效果,在KTAU上相对提高13.95%。

详情
AI中文摘要

演唱质量评估(SQA)在实际多媒体应用和音乐人工智能系统中变得越来越重要,但现有研究主要集中在短演唱片段,对全长歌曲的评估还不够。全长歌曲SQA需要对演唱质量在不同音频段的变化以及这些局部变化如何影响整体演唱表现评估进行建模。此外,段级注释的稀缺使得有效监督具有挑战性。为应对这些挑战,我们提出了SongSQA,这是一个用于全长歌曲SQA的两阶段框架。第一阶段,使用预训练教师模型生成的伪标签训练段分数预测器,无需手动段注释即可进行段级演唱质量预测。第二阶段,歌曲质量聚合器将段特征和预测的段分数集成到统一的段嵌入中,并使用可学习的歌曲嵌入和自注意力来捕捉段级演唱表现与整体歌曲质量之间的联系。实验结果证明了SongSQA对全长歌曲SQA的有效性,在KTAU上比最强基线相对提高了13.95%,同时在所有数据集上持续改进其他评估指标。

英文摘要

Singing Quality Assessment (SQA) has become increasingly important for practical multimedia applications and Music AI systems, yet existing studies predominantly focus on short singing clips and remain insufficient for full-length songs. Unlike clip-level assessment, full-length song SQA requires modeling how singing quality varies across different audio segments and how these local variations influence the overall evaluation of vocal performance. Moreover, the scarcity of segment-level annotations makes effective supervision challenging, as directly assigning a single overall score label to every segment tends to treat different segment qualities as equivalent. To address these challenges, we propose SongSQA, a two-stage framework for full-length song SQA. In the first stage, a Segment Score Predictor is trained with pseudo labels generated by a pre-trained teacher model, enabling segment-level singing quality prediction without requiring manual segment annotations. In the second stage, a Song Quality Aggregator integrates segment features and predicted segment scores into unified segment embeddings, and employs a learnable song embedding together with self-attention to capture the connection between segment-level vocal performance and overall song quality. In this way, SongSQA dynamically aggregates critical quality cues across the song to produce a holistic quality prediction, while also generating a temporal segment-level quality curve. Experimental results demonstrate the effectiveness of SongSQA for full-length song SQA, achieving up to a 13.95% relative improvement in KTAU over the strongest baseline, while consistently improving other evaluation metrics across all datasets.

URL PDF HTML 收藏
2607.09796 2026-07-21 cs.LG 版本更新

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels

噪声偏好标签下无元数据的元重加权直接偏好优化

Hua Qu, Yifan Li, Xiaodong Yuan

机构 * Xi’an Jiaotong University(西安交通大学)

AI总结 研究针对DPO性能依赖偏好数据质量问题,提出双层优化框架、无任务元知识驱动方法及结合中心差分近似与LoRA微调的可扩展训练方案,经实验验证该方法能在不同噪声率下提升训练性能。

Comments 36 pages, including appendices. Revised version with updated theoretical analysis, supplementary material, figures and improved table formatting

详情
AI中文摘要

直接偏好优化(DPO)已成为使大语言模型(LLMs)与人类偏好对齐的重要方法,因其无需显式奖励建模和强化学习优化。但其性能严重依赖偏好数据质量,现实中噪声偏好数据会削弱对齐性能。为此提出双层优化框架,在一定假设下可恢复干净数据下的DPO最优解。还推导了非对称标签翻转噪声下可学习加权函数的先验形式。考虑到高质量元数据难获取,提出无任务元知识驱动方法,即便无元数据也能元学习。结合中心差分近似与LoRA微调降低LLM元学习中高阶梯度的高成本,开发可扩展训练方案。在TL;DR摘要和Anthropic HH单轮对话实验表明,该方法在不同噪声率下比多个DPO基线提高了训练性能。

英文摘要

Direct Preference Optimization (DPO) has become an important method for aligning large language models (LLMs) with human preferences because it removes the need for explicit reward modeling and reinforcement learning. However, its performance depends heavily on the quality of preference data, and noisy preference data in real-world settings can weaken alignment performance. To address this issue, we propose a bilevel optimization framework and prove, under some idealized conditions, that this framework can recover the DPO optimum under clean data. We further derive a prior form for the learnable weighting function under label-flipping noise. Considering that high-quality metadata may be difficult to obtain, we propose a prompt augmentation consistency method that enables meta-learning even when metadata is completely unavailable. To reduce the high cost of higher-order gradients in LLM meta-learning, we combine central-difference approximation with LoRA fine-tuning and develop a scalable training scheme. Experiments on TL;DR summarization and Anthropic Helpful and Harmless dialogue show that the proposed method improves alignment performance over multiple DPO baselines under different noise rates.

URL PDF HTML 收藏
2604.01700 2026-07-21 cs.CV cs.MM 版本更新

Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation

视频扩散模型能否预测过去帧?双向循环一致性用于可逆插值

Lingyu Liu, Yaxiong Wang, Li Zhu, Zhedong Zheng

机构 * School of Software Engineering, Xi'an Jiaotong University(西安交通大学软件工程学院) School of Computer and Information Science, Hefei University of Technology(合肥工业大学计算机与信息科学学院) Faculty of Science and Technology, University of Macau(澳门大学科技学院)

AI总结 本文提出双向循环一致性框架,通过双向生成轨迹对称性提升视频插值的时序一致性,实验表明在37帧和73帧任务中取得最佳性能。

详情
AI中文摘要

视频帧插值旨在在给定端点之间合成逼真中间帧,同时遵循特定运动语义。尽管生成模型在视觉保真度上有所改进,但其主要以单向方式工作,缺乏自我验证时间一致性机制。受自监督学习中时间循环一致性的启发,我们提出了一种新的双向框架,强制正向和反向生成轨迹的对称性。我们的方法引入可学习的方向标记,以显式地将共享主干网络条件化为时间方向,使模型能够在单一统一架构中联合优化正向合成和反向重建。这种循环一致的监督作用作为一种强大的正则化器,确保生成的运动路径在逻辑上是可逆的。此外,我们采用课程学习策略,从短序列逐步训练模型,稳定不同持续时间的动力学。关键的是,我们的循环约束仅在训练期间应用;推理需要单次前向传递,保持基础模型的高效率。广泛的实验表明,我们的方法在37帧和73帧任务中在成像质量、运动平滑度和动态控制方面均达到最佳性能,优于强基线,且无额外计算开销。

英文摘要

Video frame interpolation aims to synthesize realistic intermediate frames between given endpoints while adhering to specific motion semantics. While recent generative models have improved visual fidelity, they predominantly operate in a unidirectional manner, lacking mechanisms to self-verify temporal consistency. This often leads to motion drift, directional ambiguity, and boundary misalignment, especially in long-range sequences. Inspired by the principle of temporal cycle-consistency in self-supervised learning, we propose a novel bidirectional framework that enforces symmetry between forward and backward generation trajectories. Our approach introduces learnable directional tokens to explicitly condition a shared backbone on temporal orientation, enabling the model to jointly optimize forward synthesis and backward reconstruction within a single unified architecture. This cycle-consistent supervision acts as a powerful regularizer, ensuring that generated motion paths are logically reversible. Furthermore, we employ a curriculum learning strategy that progressively trains the model from short to long sequences, stabilizing dynamics across varying durations. Crucially, our cyclic constraints are applied only during training; inference requires a single forward pass, maintaining the high efficiency of the base model. Extensive experiments show that our method achieves state-of-the-art performance in imaging quality, motion smoothness, and dynamic control on both 37-frame and 73-frame tasks, outperforming strong baselines while incurring no additional computational overhead.

URL PDF HTML 收藏
2607.15736 2026-07-20 cs.CL 新提交

Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning

更好的开始,更好的结束:用于压缩推理的自引导迭代自推理蒸馏

Leichao Dong, Dongxu Zhang, Yiding Sun, Qirui Wang, Yuhan Wang, Lin Chen, Jihua Zhu

机构 * Xi’an Jiaotong University(西安交通大学) Peking University(北京大学)

AI总结 研究大型推理模型冗余计算问题,提出BIRD两阶段自推理蒸馏方法,先在简洁指令下采样简洁解并学习,再用简洁自教师进行策略内蒸馏,在Qwen3系列模型上提升精度并降低响应长度,凸显前缀支持对高效推理蒸馏的关键作用。

详情
AI中文摘要

大型推理模型常通过长思维链解决问题,但大量计算耗费在冗余推导等上。现有策略内自蒸馏方法存在初始化瓶颈。本文提出BIRD(自引导迭代自推理蒸馏),一种两阶段自推理蒸馏方法。首先在简洁指令下从基础模型采样简洁解,保留正确答案轨迹并执行轻量级提示切换SFT步骤。然后从这个预热模型开始,使用简洁自教师进行策略内反向KL蒸馏。在Qwen3系列模型上,BIRD在MATH - 500和AIME基准测试中比提示和冷启动策略内蒸馏实现了更强的精度 - 效率权衡。如在Qwen3 - 8B上,提高了MATH - 500精度,降低了平均响应长度。结果凸显前缀支持是高效推理蒸馏的核心因素。

英文摘要

Large reasoning models often solve problems through long chain-of-thought (CoT) traces, yet much of this computation is spent on redundant derivations, repeated self-verification, and detours that do not improve the final answer. Existing on-policy self-distillation methods reduce this cost by matching a student model to a concise copy of itself on prefixes sampled from the student's own rollouts. We show that this objective has an initialization bottleneck. Since supervision is applied only to visited prefixes, training from a verbose base model places the KL loss on contexts that are often noisy, redundant, or already off track. In such regions, a concise teacher can provide only local corrections, while the student continues to explore trajectories that an efficient reasoner should avoid. In this paper, we propose BIRD(Bootstrapped Iterative Self-Reasoning Distillation), a two-stage self-reasoning distillation method that improves the rollout distribution before on-policy training. BIRD first samples concise solutions from the base model under a brevity instruction, keeps only answer-correct traces, and performs a lightweight prompt-switch SFT step. The traces are generated with the brevity instruction but learned under the original task prompt, turning instruction-induced conciseness into a default reasoning behavior. Starting from this warm model, BIRD then applies on-policy reverse-KL distillation with a concise self-teacher, now on cleaner and more informative prefixes. Across Qwen3 series models, BIRD achieves a stronger accuracy-efficiency trade-off than prompting and cold-start on-policy distillation on MATH-500 and AIME benchmarks. On Qwen3-8B, it improves MATH-500 accuracy from 86.2% to 92.0% while reducing the average response length from 3,099 to 1,115 tokens. These results highlight prefix support as a central factor in efficient reasoning distillation.

URL PDF HTML 收藏
2607.12273 2026-07-20 cs.SE cs.AI cs.CL 版本更新

Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

Code-MUE:通过基于执行的语义交互图测量代码语言模型的不确定性

Xiaoning Ren, Yinxing Xue, Lei Ma, Yuheng Huang

机构 * Xi’an Jiaotong University(西安交通大学) Institute of AI for Industries, Chinese Academy of Sciences(中国科学院人工智能产业研究院) The University of Tokyo(东京大学) University of Alberta(阿尔伯塔大学)

AI总结 研究针对代码语言模型内在随机性带来的风险,引入纯黑盒框架Code-MUE,通过基于执行的语义交互图测量不确定性,经大规模实证研究验证其与功能正确性强负相关,优于基线,可实现风险检测和选择性预测。

Comments To appear at The ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA) 2026

详情
AI中文摘要

随着代码大语言模型(LLMs)成为现代软件工程的核心,其内在的随机性带来了重大现实风险,小错误也可能导致严重后果。现有不确定性估计方法存在差距,白盒和灰盒技术不适用于闭源模型,标准黑盒文本指标无法捕捉代码独特脆弱性。为此引入Code-MUE,一个通过基于执行的语义交互图测量不确定性的纯黑盒框架。它基于可观察运行时行为计算解空间的冯·诺依曼熵来量化全局语义多样性。大规模实证研究表明,Code-MUE与功能正确性呈强负相关,显著优于基于词汇和嵌入的基线,能在实际工作流程中实现强大的风险检测和选择性预测。

英文摘要

As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, where even minor errors can lead to severe functional, security, or safety consequences. Reliable automation, therefore, demands the ability to distinguish between confident, well-supported predictions and stochastic guessing. However, existing uncertainty estimation methods face a critical gap: white and grey-box techniques are often inapplicable to closed-source models, while standard "black-box" text metrics fail to capture the unique fragility of code, where syntactic variation does not always imply semantic divergence. To bridge this syntax-semantics gap, we introduce Code-MUE, a purely black-box framework that measures uncertainty through execution-based Semantic Interaction Graphs. Different from prior approaches that rely on superficial textual similarity, Code-MUE grounds uncertainty in observable runtime behavior, calculating the Von Neumann entropy of the solution space to quantify global semantic diversity. A large-scale empirical study across eight state-of-the-art LLMs demonstrates that Code-MUE achieves a strong negative correlation with functional correctness (Spearman's correlation up to -0.98), significantly outperforming lexical and embedding-based baselines while enabling robust risk detection and selective prediction in practical workflows.

URL PDF HTML 收藏
2601.12222 2026-07-20 cs.SD cs.MM eess.AS 版本更新

Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling

基于多茎注意力和分层不确定性建模的歌曲美学评估

Yishan Lv, Jing Luo, Boyuan Ju, Yang Zhang, Xinda Wu, Bo Yuan, Xinyu Yang

机构 * School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an, China(计算机科学与技术学院,西安交通大学,西安,中国) Central Media Technology Institute, Huawei(中央媒体技术研究院,华为)

AI总结 针对音乐生成人工智能带来的歌曲美学评估需求,提出含多茎注意力融合与分层粒度感知区间聚合模块的评估框架,在两个数据集上评估并与两个SOTA模型比较,该方法在多维歌曲美学评估中性能更强。

Comments Accepted to the 27th International Society for Music Information Retrieval Conference (ISMIR 2026)

详情
AI中文摘要

音乐生成人工智能正在迅速扩展音乐内容,因此需要自动化的歌曲美学评估。然而,现有研究大多集中在语音、音频或演唱质量上,歌曲美学研究不足。此外,传统方法通常直接预测精确的平均意见得分(MOS)值,难以捕捉歌曲美学评估中人类感知的细微差别。本文提出了一个面向歌曲的美学评估框架,具有两个新颖的模块:多茎注意力融合(MSAF)在混合人声和混合伴奏对之间建立双向交叉注意力,融合它们以捕捉复杂的音乐特征;分层粒度感知区间聚合(HiGIA)学习多粒度得分概率分布,将它们聚合到一个得分区间,并在区间内应用回归以产生最终得分。我们在两个全长歌曲数据集上进行了评估:SongEval数据集(人工智能生成)和一个内部美学数据集(人类创作),并与两个最先进的(SOTA)模型进行了比较。结果表明,所提出的方法在多维歌曲美学评估中取得了更强的性能。推理代码和检查点可在这个https URL上公开获得。

英文摘要

Music generative artificial intelligence (AI) is rapidly expanding music content, necessitating automated song aesthetics evaluation. However, existing studies largely focus on speech, audio or singing quality, leaving song aesthetics underexplored. Moreover, conventional approaches often predict a precise Mean Opinion Score (MOS) value directly, which struggles to capture the nuances of human perception in song aesthetics evaluation. This paper proposes a song-oriented aesthetics evaluation framework, featuring two novel modules: 1) Multi-Stem Attention Fusion (MSAF) builds bidirectional cross-attention between mixture-vocal and mixture-accompaniment pairs, fusing them to capture complex musical features; 2) Hierarchical Granularity-Aware Interval Aggregation (HiGIA) learns multi-granularity score probability distributions, aggregates them into a score interval, and applies a regression within the interval to produce the final score. We evaluated on two datasets of full-length songs: SongEval dataset (AI-generated) and an internal aesthetics dataset (human-created), and compared with two state-of-the-art (SOTA) models. Results show that the proposed method achieves stronger performance for multi-dimensional song aesthetics evaluation. The inference code and checkpoint are publicly available at https://github.com/yisan33/song-aesthetics-evaluation.

URL PDF HTML 收藏
2607.14976 2026-07-17 cs.CV 新提交

From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting

从带草稿到无草稿:通过特权蒸馏和快速植入实现一步视频对象移除

Zizhao Chen, Ping Wei, Guang Dai, Jingdong Wang, Mengmeng Wang

机构 * Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人工智能与机器人研究所) SGIT AI Lab, State Grid Corporation of China(国家电网公司SGIT人工智能实验室) Zhejiang University of Technology(浙江工业大学) Baidu(百度)

AI总结 研究视频对象移除问题,提出D2DF框架,通过特权蒸馏和自引导快速植入模块,将教师模型多步细化能力提炼到学生模型,实现一步视频对象移除,在质量和效率上优于传统及多步生成方法。

Comments Accepted by ECCV 2026

详情
AI中文摘要

视频对象移除是视频编辑中一项基本但具有挑战性的任务。尽管近期有进展,但现有方法分为两类。传统方法常引入明显伪影且结果不自然,基于扩散的方法视觉效果好但需多步去噪,实用性受限。我们提出D2DF框架,将把粗糙草稿转化为精细视频的能力提炼到一步视频生成模型中。训练教师模型将低质量移除结果细化为高保真视频,通过PPCD将此能力提炼到学生模型。引入SGFP模块消除对草稿的依赖,实现完全无草稿的一步模型。实验表明,有草稿和无草稿版本在多个指标上均达最优,在质量和效率上超越传统及多步生成方法,单视频去噪仅需约1秒。

英文摘要

Video object removal is a fundamental yet challenging task in video editing. Despite recent progress, existing methods typically fall into two categories. Traditional approaches based on optical flow or attention mechanisms often introduce noticeable artifacts and yield unnatural results. In contrast, diffusion-based methods improve visual realism but demand multiple denoising steps, limiting their practicality. To address these issues, we propose From-Draft-to-Draft-Free (D2DF), a framework that distills the ability of transforming coarse drafts into refined videos into a one-step video generation model. Within D2DF, a teacher model is trained to refine low-quality removal results ("drafts") into high-fidelity videos by multiple steps. Then, through Prior-Privileged Consistency Distillation (PPCD), we distill this capability into a student model that performs one-step removal conditioned on the draft. To eliminate draft dependency, we introduce a Self-Guided Fast Planting (SGFP) module based on our Temporal Masked Transformer that autonomously generates scene-consistent pseudo-drafts in latent space, enabling a fully draft-free one-step model. Extensive experiments show that both draft-conditioned and draft-free versions achieve state-of-the-art performance on multiple metrics, surpassing traditional and multi-step generative methods in both quality and efficiency. The denoising process for a single video takes only about 1 second.

URL PDF HTML 收藏
2607.14974 2026-07-17 cs.CV cs.CR 新提交

On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline

关于成功与简洁:对可转移视觉语言攻击管道的再审视

Yuchen Ren, Zhengyu Zhao, Chenhao Lin, Bo Yang, Chao Shen

机构 * School of Cyber Science and Engineering, Xi’an Jiaotong University(西安交通大学网络空间安全学院) State Key Laboratory of Mathematical Engineering and Advanced Computing, Information Engineering University(信息工程大学数学工程与先进计算国家重点实验室)

AI总结 研究视觉语言预训练模型对抗攻击,指出现有复杂攻击管道可简化。提出SimVLA管道解决问题,经实验验证在可转移性和效率上优于基线,强调利用领域知识的重要性,为未来扩展提供简单有效主干。

Comments Accepted for publication in IEEE Transactions on Information Forensics and Security (TIFS)

详情
AI中文摘要

视觉语言预训练模型(VLPMs)容易受到对抗攻击。近期对VLPMs的可转移攻击遵循复杂损失函数或多阶段文本/图像攻击的通用管道。本文表明复杂攻击管道可更简单且更成功。识别出由不当跨模态交互和过多操作导致的三个被忽视问题,提出简单视觉语言攻击(SimVLA)管道,提高了可转移性和效率。在四个数据集和三个下游任务上实验验证了其优越性,如在Flickr30k文本图像检索数据集上,SimVLA在R@1可转移性上比SOTA基线高出8.01%-14.71%,同时仅消耗约35.73%的时间和46.26%的最大VRAM。突出了利用领域知识的重要性,盲目追求复杂操作可能有害,希望SimVLA能成为未来扩展的简单有效主干。代码可通过链接获取。

英文摘要

Vision-Language Pre-training Models (VLPMs) are known to be vulnerable to adversarial attacks. Recent transferable attacks on VLPMs have followed a common pipeline with complicated loss functions or multi-stage text/image attacks. However, in this paper, we demonstrate that such a sophisticated attack pipeline can be simpler yet more successful. Specifically, we identify three previously overlooked issues caused by inappropriate cross-modal interactions and excessive operations. To address them, we propose the Simple Vision-Language Attack (SimVLA) pipeline, which observably improves transferability and efficiency. Experiments on four datasets and three downstream tasks validate the superiority of our pipeline. For instance, on Flickr30k text-image retrieval dataset, our SimVLA outperforms the SOTA baseline in R@1 transferability by 8.01\%-14.71\%, while consuming only about 35.73\% of the time and 46.26\% of the max VRAM. Overall, the superiority of our SimVLA highlights the importance of leveraging domain knowledge (e.g., our proposed cross-modal word identification), while blindly pursuing intricate operations (e.g, complex loss functions and redundant multi-stage designs) may even be harmful. We hope our SimVLA can serve as a simple yet effective backbone for future extensions. Code is available at https://github.com/RYC-98/SimVLA.

URL PDF HTML 收藏
2607.13731 2026-07-16 cs.LG stat.ML 新提交

DAGR: State-Conditioned Goal Representations via Difference-Aware Goal Cross-Attention

DAGR:通过差异感知目标交叉注意力实现的状态条件目标表示

Xing Lei, Wenyan Yang, Xuetao Zhang, Donglin Wang

机构 * Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人工智能与机器人研究所) Department of Electrical Engineering and Automation, Aalto University(阿尔托大学电气工程与自动化系) School of Engineering, Westlake University(西湖大学工学院)

AI总结 研究目标条件强化学习中目标编码问题,提出DAGR方法,通过多尺度门控交叉注意力将静态嵌入细化为状态条件嵌入,在OGBench上改善导航,在操纵和拼图任务中有不同表现,是一种结构化细化。

详情
AI中文摘要

目标条件强化学习取决于目标的编码方式。对比、度量、时间距离和信息理论编码器在目标上有所不同,但都有一个共同特点,即都不考虑当前状态。这种与状态无关的嵌入无法标记目标中仍需采取行动的部分,策略必须通过反转两个编码器来恢复该线索。我们提出了DAGR,它通过多尺度门控交叉注意力将任何后期融合编码器的静态嵌入细化为状态条件嵌入。近恒等门控残差保留了基础表示。差异感知目标交叉注意力然后使用每个令牌的状态-目标差异图来偏向注意力分数。在OGBench上,DAGR改善了导航。我们的消融实验将收益追溯到门控残差,而不是命名该方法的差异偏差。在操纵和拼图任务中,它与基础模型匹配或低于基础模型。DAGR是一种结构化细化,而不是普遍改进。

英文摘要

Goal-conditioned reinforcement learning hinges on how the goal is encoded. Contrastive, metric, temporal-distance, and information-theoretic encoders differ in objective. They still share one trait. None of them sees the current state. Such a state-independent embedding cannot mark which part of the goal still needs action. The policy must then recover that cue by inverting both encoders. We propose DAGR. It refines the static embedding of any late-fusion encoder into a state-conditioned one through multi-scale gated cross-attention. A near-identity gated residual preserves the base representation. Difference-aware Goal Cross-Attention then biases the attention scores using a per-token state-goal discrepancy map. On OGBench, DAGR improves navigation. Our ablations trace the gain to the gated residual, not to the difference bias that names the method. On manipulation and puzzle tasks it matches or falls below the base. DAGR is a structured refinement, not a universal improvement.

URL PDF HTML 收藏
2607.13563 2026-07-16 cs.CV 新提交

Nexus: Native Mesh Generation with Diffusion

Nexus:基于扩散的原生网格生成

Hanxiao Wang, Ying-Tian Liu, Yuan-Chen Guo, Qi-Yuan Feng, Zi-Xin Zou, Ding Liang, Biao Zhang, Yan-Pei Cao

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统管理与控制国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) VAST(无(暂未找到合适中文名,VAST可音译为“瓦斯特”,但这并非一个正式的中文机构名,所以保留英文)) Tsinghua University(清华大学) Xi’an Jiaotong University(西安交通大学)

AI总结 研究针对高质量三角形网格生成问题,提出Nexus扩散方法,通过解耦顶点与拓扑生成实现整体网格生成,经实验验证其性能优于现有基线,有效克服顺序网格建模局限且获从业者青睐。

详情
AI中文摘要

生成高质量三角形网格对电影、游戏和交互式3D应用至关重要。主流方法依赖网格序列化和自回归过程,在有效推理方面存在困难且对误差积累敏感。本文提出Nexus,一种通过解耦顶点和拓扑生成实现整体网格生成的扩散方法。先将网格顶点视为八叉树组织的稀疏体素,用扩散模型从粗到细生成顶点;为拓扑建模提出时空间隔,将任意边和面拓扑编码为连续顶点嵌入,再用扩散模型在生成顶点上生成连续嵌入。在Objaverse和Toys4K数据集及自然图像上的大量实验表明,该方法优于现有自回归和两阶段基线,有效规避顺序网格建模固有局限,3D从业者的盲测显示对其结果有强烈感知偏好。

英文摘要

Generating high-quality triangle meshes is essential for film, gaming, and interactive 3D applications. Mainstream methods rely on mesh serialization and autoregressive processes, which stuggles in effective inference and is sensitive to error accumulation. In this paper, we present Nexus, a diffusion method that achieves holistic mesh generation via decoupled vertex and topology generation. First, we view mesh vertices as sparse voxels organized as an octree and adopt a diffusion model to generate the vertices in a coarse-to-fine manner. Second, for topology modeling, we propose Spacetime Interval, as an extension of Spacetime Distance to encode arbitrary edge and face topology into continuous per-vertex embeddings. It allows for a global and efficient recovery of complex topology. We then employ a diffusion model to generate the continuous embeddings on the generated vertices. Extensive experiments on the Objaverse and Toys4K datasets and in-the-wild images demonstrate that our method outperforms state-of-the-art autoregressive and two-stage baselines, effectively circumventing the inherent limitations of sequential mesh modeling. A blind user study from 3D practitioners confirms strong perceptual preference for our results.

URL PDF HTML 收藏
2607.12992 2026-07-15 cs.RO 新提交

ChunkFlow: Towards Continuity-Consistent Chunked Policy Learning

ChunkFlow:迈向连续性一致的分块策略学习

Zhao Yang, Yinan Shi, Mingyuan Yao, Wenyao Xue, Yawei Jueluo, Longjun Liu

机构 * Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人工智能与机器人研究所) Jiangsu Cytoderm Intelligent Technology Co., Ltd.(江苏希迪姆智能科技有限公司)

AI总结 研究视觉语言动作模型分块策略的边界抖动问题,提出ChunkFlow框架划分块区域,执行时用确定性重叠混合,训练时用多种损失,通过实验验证该框架在低延迟推理下能改善成功稳定性权衡。

详情
AI中文摘要

视觉语言动作(VLA)模型越来越多地采用分块动作头来满足实时约束,但这会引入边界抖动,连续块之间的重叠区域常产生不一致预测,降低时间连贯性和任务成功率。现有方法如推理时混合仅重新加权不匹配提议而不纠正根本错误。我们提出ChunkFlow,一种用于分块策略的感知接缝训练和执行框架,将块结构与边界执行对齐。它划分区域,执行时应用确定性重叠混合,用接缝及一阶和二阶连续性损失训练原始预测。实验表明其在低延迟推理下改善了成功稳定性权衡。

英文摘要

Vision-language action (VLA) models increasingly adopt chunked action heads to satisfy real-time constraints; however, this introduces boundary jitter: overlapping regions between consecutive chunks often yield inconsistent predictions, degrading temporal coherence and the task success rate. Existing methods, such as inference-time blending, merely reweight mismatched proposals without correcting underlying errors, leading to residual accumulation under biased or noisy histories. We propose ChunkFlow, a seam-aware training-and-execution framework for chunked policies that aligns chunk structure with boundary execution. It partitions each chunk into frozen, editable, and future zones, applies deterministic overlap blending at execution, and trains raw predictions with seam and first- and second-order continuity losses. History corruption and scheduled sampling improve robustness to executed-history errors, while an AWAC fine-tuning stage adapts the policy without removing these structural regularizers. Under mild smoothness assumptions, pre-blending seam discrepancies provably decay with increasing overlap. Experiments on CALVIN, LIBERO, and real robots show an improved success-stability trade-off with low-latency inference. Project page: https://cytoderm-ai.github.io/chunkflow.

URL PDF HTML 收藏
2607.12786 2026-07-15 cs.CV 新提交

CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models

CoRe:视觉语言模型中跨图像比较推理的综合框架

Lin Peng, Cong Wan, Zeyu Guo, SongLin Dong, Yihong Gong

机构 * Xi’an Jiaotong University(西安交通大学)

AI总结 针对视觉语言模型跨图像比较推理难题,提出CoRe框架,含自动构建的训练集CoRe-20K、结构化奖励框架TriSR及基准CoRe-Bench,实验显示其在CoRe-Bench上大幅超越现有模型,在标准基准上也有竞争力。

Comments Accepted by ACMMM2026

详情
AI中文摘要

跨图像比较推理对视觉语言模型(VLM)来说仍然具有挑战性,特别是在正确预测需要细粒度属性基础和全局一致推理时。我们提出了CoRe,一个针对此问题的统一框架。CoRe包括:通过多专家协作管道从结构化视觉元数据自动构建的大规模基于三元组的训练集CoRe-20K;在GRPO优化下联合监督属性基础、判断对齐和三元组一致性的结构化奖励框架TriSR;以及首个专门用于细粒度跨图像比较推理的基准CoRe-Bench。实验表明,CoRe在CoRe-Bench上显著优于现有VLM,在标准多模态基准上也具有竞争力,部分准确率比最强基线提高了28.2个百分点。

英文摘要

Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally consistent reasoning. We present CoRe, a unified framework for this problem. CoRe includes: (i) CoRe-20K, a large-scale triplet-based training set automatically constructed from structured visual metadata through a multi-expert collaborative pipeline, covering counting, depth, distance, and spatial relations; (ii) TriSR, a structured reward framework that jointly supervises attribute grounding, judgment alignment, and triplet consistency under GRPO optimization; and (iii) CoRe-Bench, the first benchmark dedicated to fine-grained cross-image comparative reasoning. Experiments show that CoRe substantially outperforms existing VLMs on CoRe-Bench while remaining competitive on standard multimodal benchmarks, achieving a 28.2-point gain in partial accuracy over the strongest baseline.

URL PDF HTML 收藏
2607.11557 2026-07-14 cs.CV 新提交

Single-Teacher View Augmentation: Enhancing Knowledge Distillation with Student-Guided Perturbations

单教师视图增强:通过学生引导的扰动增强知识蒸馏

Xuyi Yu, Yaohua Liu, Ziming Song, Yinghai Zhao, Huipeng Zhang, Kuizhi Mei

机构 * State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人工智能与机器人研究所人机混合增强智能技术国家重点实验室) Guangdong Institute of Intelligence Science and Technology(广东省智能科学与技术研究院) Institute of Collaborative Innovation, University of Macau(澳门大学协同创新研究院) Beijing Huahang Institute of Radio Measurement(北京华航无线电测量研究所)

AI总结 研究针对知识蒸馏中单一教师视角限制监督信号多样性的问题,提出SAKD框架,利用学生演变特征动态生成扰动视图,实现单阶段训练,实验证明该方法在减少参数和无需预训练的情况下,准确率优于随机扰动方法且与两阶段方法相当。

详情
AI中文摘要

知识蒸馏通常依赖单一教师的固定视角,限制了监督信号的多样性。多教师蒸馏虽能解决此问题,但计算和存储成本过高。为平衡效率与多样性,近期研究聚焦于从单一教师生成虚拟视图。现有方法存在权衡:随机扰动方法高效但缺乏可控多样性,结构化增强方法需多阶段训练且参数线性增长。我们提出Shift-Augmented Knowledge Distillation(SAKD)框架,利用学生不断演变的特征作为扰动生成的动态条件,实现单阶段训练并通过无参数循环移位产生自适应、多样的视图。在CIFAR-100和ImageNet上的大量实验表明,SAKD始终优于随机扰动方法,且在参数显著减少并消除预训练要求的情况下,达到与两阶段方法相当的准确率。

英文摘要

Knowledge distillation (KD) typically relies on the fixed perspective of a single teacher, limiting the diversity of supervisory signals. While multi-teacher distillation addresses this by aggregating knowledge from multiple models, it incurs prohibitive computational and storage costs. To balance efficiency and diversity, recent research has focused on generating virtual views from a single teacher. However, existing methods face a trade-off: random perturbation approaches offer efficiency but lack controlled diversity, while structured augmentation methods require multi-stage training and incur linear parameter growth. We observe that this trade-off stems from a common design choice: using the teacher's strong but static features to generate views. Instead, we propose Shift-Augmented Knowledge Distillation (SAKD), a simple yet effective framework that leverages the student's evolving features as a dynamic condition for perturbation generation. This shift in perspective enables single-stage training while producing adaptive, diverse views through a parameter-free cyclic shift. Extensive experiments on CIFAR-100 and ImageNet demonstrate that SAKD consistently outperforms random perturbation methods and achieves accuracy on par with two-stage approaches, while using significantly fewer parameters and eliminating pre-training requirements.

URL PDF HTML 收藏
2607.10792 2026-07-14 cs.CV 新提交

MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction

MAC-Splat:用于高保真稀疏视图重建的多属性一致性

Jinqian Yang, Yichen Wu, Wanhua Li, Haokun Lin, Renzhen Wang, Xiangchu Feng, Xixi Jia

机构 * Xidian University(西安电子科技大学) Harvard University(哈佛大学) Nanyang Technological University(南洋理工大学) City University of Hong Kong(香港城市大学) Xi’an Jiaotong University(西安交通大学)

AI总结 针对稀疏视图重建中现有方法存在几何伪影的问题,提出MAC-Splat训练框架,利用MASt3R和DINOv3获取2D对应关系并定义MAC损失,联合正则化3D属性,实验证明该方法能有效解决不适定的稀疏视图重建问题,性能优于基线。

Comments Accepted to the European Conference on Computer Vision (ECCV 2026)

详情
AI中文摘要

从稀疏视图重建高保真3D场景一直是可推广神经渲染中的核心问题。现有的可推广3D高斯喷溅(3DGS)方法在稀疏视图设置中常出现几何伪影,因为仅基于2D光度损失的监督无法解决深度和对应模糊性。为此提出MAC-Splat,一个围绕直接3D一致性监督构建的训练框架。它基于MASt3R几何主干和冻结的DINOv3编码器获取语义信息丰富的2D对应关系,以此定义多属性一致性(MAC)损失,联合正则化匹配高斯的3D属性。实验表明MAC-Splat优于强基线,尤其在不同重叠情况下有显著提升,有效解决不适定的稀疏视图重建问题。

英文摘要

Reconstructing high-fidelity 3D scenes from sparse-views remains a central problem in generalizable neural rendering. Existing generalizable 3D Gaussian Splatting (3DGS) methods often exhibit geometric artifacts in sparse-view settings, since supervision based solely on 2D photometric losses cannot resolve depth and correspondence ambiguities. To address this issue, we propose MAC-Splat, a training framework built around direct 3D consistency supervision. MAC-Splat builds on the MASt3R geometric backbone and a frozen DINOv3 encoder to obtain semantically informed 2D correspondences, which serve as geometric anchors for 3D supervision. Using these anchors, we define the Multi-Attribute Consistency (MAC) loss. This objective jointly regularizes the 3D attributes of matched Gaussians, including their position, shape, and appearance, by enforcing agreement in a common world coordinate frame. The formulation is robust to outliers and respects the geometry of covariance matrices, which leads to stable training under sparse-view conditions. Experiments on ScanNet++ show that MAC-Splat outperforms strong baselines, with particularly large gains under different overlap regimes. In particular, it improves average PSNR over Splatt3R by more than 4.5 dB, reduces LPIPS, and maintains performance as the camera pose gap increases. These results indicate that a direct, multi-attribute 3D consistency objective, when combined with high-quality correspondences, is effective for addressing the ill-posed sparse-view reconstruction problem.

URL PDF HTML 收藏
2607.10296 2026-07-14 cs.AI cs.CL 新提交

SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models

SPARK:大语言模型中基于敏感性的潜在推理状态分析与引导

Dongxu Zhang, Yiding Sun, Zihao Guo, Xiangyang Yang, Kai Tang, Lin Chen, Cheng Tan, Jihua Zhu

机构 * Xi’an Jiaotong University(西安交通大学) Peking University(北京大学) Tencent(腾讯)

AI总结 研究大语言模型推理失败问题,提出SPARK方法,利用隐藏状态响应诊断推理状态并引导测试时干预,通过长度控制敏感性等手段,在实验中提升了Qwen3系列模型性能,证明敏感性对推理失败诊断及干预的作用。

详情
AI中文摘要

大语言模型中的推理失败通常从最终答案评估,但错误答案无法揭示失败原因。现有方法多在输出层面操作,通用激活引导方法未诊断哪些示例需干预。本文介绍SPARK,利用隐藏状态响应诊断模型是否进入有效推理状态并引导轻量级测试时引导。原始隐藏状态敏感性受提示长度强烈混淆,SPARK用长度控制敏感性分离输入规模效应与残余推理激活,结合该信号与跨层协调选择推理活跃锚点和未充分激活的难示例。通过实验,该方法持续提升Qwen3系列模型性能,表明敏感性不仅可作为推理失败的诊断信号,还可作为针对性测试时干预的实用指南。

英文摘要

Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may reflect missing capability, an unstable reasoning trajectory, or a failure to activate a reasoning state that is already available in the frozen model. Existing prompting and benchmark-based evaluation methods mostly operate at the output level, while generic activation-steering methods typically apply global directions without diagnosing which examples require intervention. In this paper, we introduce SPARK, which uses hidden-state response to diagnose whether a model internally enters an effective reasoning state and to guide lightweight test-time steering. The key observation is that raw hidden-state susceptibility is strongly confounded by prompt length, especially in programmatic and algorithmic reasoning where harder serialized instances naturally become longer. SPARK therefore uses length-controlled susceptibility to separate input-scale effects from residual reasoning activation, and combines this signal with cross-layer coordination to select reasoning-active anchors and under-activated hard examples. We use FRONTIER-4.5K as a controlled programmatic reasoning suite for latent profiling and difficulty-aware analysis, and evaluate SPARK-Steering on GSM8K and MATH-500 with forward-only benchmark profiling. Our method improves Qwen3 series models consistently; on MATH-500, accuracy rises from 82.0% to 84.6% for Qwen3-4B and from 82.4% to 85.6% for Qwen3-8B. These results suggest that susceptibility can serve not only as a diagnostic signal for reasoning failures, but also as a practical guide for targeted test-time intervention.

URL PDF HTML 收藏
2607.10263 2026-07-14 cs.LG 新提交

Sharper Analysis of Single-Loop Methods for Bilevel Optimization

双层优化单环方法的更精确分析

Yubo Zhou, Jun Shu, Luo Luo, Junmin Liu, Deyu Meng, Guang Dai, Haishan Ye

机构 * School of Mathematics and Statistics, Xi’an Jiaotong University(西安交通大学数学与统计学院) School of Data Science, Fudan University(复旦大学数据科学学院) SGIT AI Lab, State Grid Corporation of China(国家电网公司SGIT人工智能实验室) Center for Intelligent Decision-Making and Machine Learning, School of Management, Xi’an Jiaotong University(西安交通大学管理学院智能决策与机器学习中心)

AI总结 研究双层优化中理论与实际单环实现的差距,利用解耦范数分析框架,改进单环近似隐式微分和迭代微分方法的收敛结果,提升收敛速率并精确匹配渐近误差下界,经实验验证理论发现。

Comments 26 pages,6 figures

详情
AI中文摘要

双层优化支撑着许多机器学习应用,如超参数优化、元学习、神经架构搜索和强化学习。基于超梯度的方法虽有显著进展,但理论保证与高效的实际单环实现之间仍存在差距。我们利用提出的解耦范数分析(DNA)框架,为单环近似隐式微分(AID)和迭代微分(ITD)方法建立了更精确的收敛结果,弥合了这一差距。对于AID,将收敛速率从\(\mathcal{O}(\kappa^6/K)\)提高到\(\mathcal{O}(\kappa^5/K)\);对于ITD,证明渐近误差为\(\mathcal{O}(\kappa^2)\),与已知下界精确匹配并改进了先前的\(\mathcal{O}(\kappa^3)\)保证。合成和实际任务的数值实验证实了我们的理论发现。

英文摘要

Bilevel optimization underpins many machine learning applications, including hyperparameter optimization, meta-learning, neural architecture search, and reinforcement learning. While hypergradient-based methods have advanced significantly, a gap persists between theoretical guarantees and practical single-loop implementations required for efficiency. We bridge this gap by establishing sharper convergence results for single-loop approximate implicit differentiation (AID) and iterative differentiation (ITD) methods, leveraging our proposed analytical framework, decoupled norm analysis (DNA). For AID, we improve the convergence rate from $\mathcal{O}(κ^6/K)$ to $\mathcal{O}(κ^5/K)$, where $κ$ is the condition number of the inner-level problem. For ITD, we prove that the asymptotic error is $\mathcal{O}(κ^2)$, exactly matching the known lower bound and improving upon the previous $\mathcal{O}(κ^3)$ guarantee. Numerical experiments on synthetic and real tasks corroborate our theoretical findings.

URL PDF HTML 收藏
2606.11637 2026-07-14 cs.AI 版本更新

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

TouchThinker: 通过大规模数据和动作感知表示将触觉常识推理扩展到开放世界

Kailin Lyu, Di Wu, Pengwei Zhang, Yuhang Zheng, Yingxin Lai, Long Xiao, Kangyi Wu, Pengna Li, Chen Gao, Lianyu Hu, Xiaobin Hu, Jie Hao, Ce Hao, Weihao Yuan, Shuicheng Yan

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) National University of Singapore(新加坡国立大学) Zhongguancun Academy(中关村学院) Xiamen University(厦门大学) Xi’an Jiaotong University(西安交通大学) Nanyang Technological University(南洋理工大学) Nanjing University(南京大学)

AI总结 提出TouchThinker框架,通过构建百万级多源触觉数据集TouchThinker-1M和动作感知建模,将触觉常识推理扩展到开放世界,在多个数据集上取得竞争性表现。

Comments 18 pages, 11 figures

详情
AI中文摘要

触觉是具身智能体理解物理世界的关键模态。尽管最近的工作已将触觉信号融入语言系统进行触觉常识推理,但由于两个关键瓶颈,将此类系统扩展到现实的开放世界环境仍然具有挑战性:(1) 当前的触觉推理数据集在格式和规模上仍然有限,为从触觉观察到物理常识的推理提供的监督不足,并阻碍了可迁移触觉常识的学习;(2) 触觉信号本质上是冗余且特定于动作的,但现有方法常常忽略这些特性,导致表示效率低下且语义表达能力有限。为了解决这些局限性,我们提出了TouchThinker,一个从数据和表示两个角度将触觉常识推理扩展到开放世界的触觉-语言框架。首先,我们构建了TouchThinker-1M,一个百万级、多源的触觉推理数据集,涵盖\textbf{415}个物体、\textbf{8}个场景和\textbf{7}种传感器类型,为开放世界泛化提供了坚实的数据基础。我们进一步引入了TouchThinker-Bench,一个具有更真实和多样化任务的开放世界基准。然后,我们提出了动作感知建模机制,以提高触觉表示效率并实现高效推理。实验结果表明,TouchThinker在多个数据集上取得了与最先进模型竞争的性能。我们的代码和数据集将在以下网址提供:this https URL。

英文摘要

Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realistic open-world settings remains challenging due to two key bottlenecks: (1) current tactile reasoning datasets remain limited in format and scale, providing insufficient supervision for reasoning from tactile observations to physical commonsense and hindering the learning of transferable tactile commonsense; (2) Tactile signals are inherently redundant and action-specific, yet existing methods often overlook these properties, resulting in inefficient representations with limited semantic expressiveness. To address these limitations, we propose TouchThinker, a tactile-language framework that scales tactile commonsense reasoning to the open world from both data and representation perspectives. First, we construct TouchThinker-1M, a million-scale, multi-source tactile reasoning dataset covering \textbf{415} objects, \textbf{8} scenarios, and \textbf{7} sensor types, providing a solid data foundation for open-world generalization. We further introduce TouchThinker-Bench, an open-world benchmark with more realistic and diverse tasks. Then, we propose action-aware modeling mechanism to improve tactile representation efficiency and enable efficient reasoning. Experimental results demonstrate that TouchThinker achieves competitive performance against state-of-the-art models across multiple datasets. Our code and dataset will be made available at: https://github.com/lvkailin0118/TouchThinker.

URL PDF HTML 收藏
2604.17889 2026-07-14 cs.CV 版本更新

AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning

AeroRAG:结构化多模态检索增强型LLM用于细粒度航空视觉推理

Junxiao Xue, Quan Deng, Tingqi Hu, Meicong Si, Xinyi Yin, Yunyun Shi, Xuecheng Wu

机构 * Research Center for Space Computing System, Zhejiang Lab, Hangzhou(杭州浙大实验室空间计算系统研究中心) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou(中国科学院大学杭州高等研究院) School of Cyber Science and Engineering, Zhengzhou University(郑州大学计算机科学与工程学院) School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)

AI总结 本文提出AeroRAG,通过结构化场景图引导的多模态检索增强生成框架,解决航空场景中视觉问答的挑战,提升对小物体、数量、位置及物体关系的推理能力。

详情
AI中文摘要

尽管多模态大语言模型(MLLMs)在多模态任务上取得了进展,但可靠的航空场景视觉问答仍具挑战性。在这些场景中,任务关键证据通常由小物体、明确数量、粗略位置及物体间关系承载,而传统密集视觉令牌表示与这些结构化语义不匹配。为此,我们提出AeroRAG,一种基于场景图的多模态检索增强生成框架,用于视觉问答。该框架首先将输入图像转换为结构化的视觉知识,包括物体类别、数量、空间位置及语义关系,然后检索与查询相关的语义片段,构建紧凑的提示以供基于文本的大型语言模型使用。与直接在密集视觉令牌上进行推理不同,我们的方法引入了感知与语言推理之间的更明确的中间接口。在AUG航空数据集和通用领域VG-150基准上的实验表明,AeroRAG在六个强大的MLLM基线中表现出一致的改进,尤其在密集航空场景和关系敏感推理中收益最大。我们进一步在VQAv2上评估该框架,验证所提出的接口仍能兼容标准视觉推理设置。这些结果表明,结构化检索是部署导向和基于事实的视觉推理系统的一种实用设计方向。

英文摘要

Despite recent progress in multimodal large language models (MLLMs), reliable visual question answering in aerial scenes remains challenging. In such scenes, task-critical evidence is often carried by small objects, explicit quantities, coarse locations, and inter-object relations, whereas conventional dense visual-token representations are not well aligned with these structured semantics. To address this interface mismatch, we propose AeroRAG, a scene-graph-guided multimodal retrieval-augmented generation framework for visual question answering. The framework first converts an input image into structured visual knowledge, including object categories, quantities, spatial locations, and semantic relations, and then retrieves query-relevant semantic chunks to construct compact prompts for a text-based large language model. Rather than relying on direct reasoning over dense visual tokens, our method introduces a more explicit intermediate interface between perception and language reasoning. Experiments on the AUG aerial dataset and the general-domain VG-150 benchmark show consistent improvements over six strong MLLM baselines, with the largest gains observed in dense aerial scenes and relation-sensitive reasoning. We further evaluate the framework on VQAv2 to verify that the proposed interface remains compatible with standard visual reasoning settings. These results suggest that structured retrieval is a practical design direction for deployment-oriented and grounded visual reasoning systems.

URL PDF HTML 收藏
2504.15650 2026-07-13 cs.CV 版本更新

AffordanceSAM: Segment Anything Once More in Affordance Grounding

可负担性语义分割模型:在可负担性基础上再次分割任何事物

Dengyang Jiang, Zanyi Wang, Hengzhuang Li, Sizhe Dang, Teli Ma, Wei Wei, Guang Dai, Lei Zhang, Harry Yang, Mengmeng Wang

机构 * SGIT AI Lab(SGIT人工智能实验室) NWPU(西北工业大学) HKUST(香港科技大学) XJTU(西安交通大学) HUST(华中科技大学) ZJUT(浙江工业大学)

AI总结 研究聚焦全监督可负担性基础,提出AffordanceSAM,通过设计适应模块和标注数据集,以三阶段训练方式扩展SAM泛化能力,在AGD20K基准上达最优性能,展现强大泛化能力。

Comments [ACM MM 2026] SAM Meets Affordance Grounding

详情
AI中文摘要

构建一个广义的可负担性基础模型以识别物体上的可操作区域对实际应用至关重要。现有训练该模型的方法分为弱监督和全监督方式。前者需要复杂训练框架设计且无辅助先验时无法推断新动作,后者虽简单但受限于有限标注数据和从头训练的组件。本研究聚焦全监督可负担性基础,提出AffordanceSAM克服其局限性,将SAM在分割中的泛化能力扩展到可负担性基础。具体设计了可负担性适应模块并策划了名为C2F - Aff的从粗到细标注数据集,以三阶段训练方式将SAM的强大性能转移到可负担性上。实验结果证实AffordanceSAM在AGD20K基准上达到了当前最优性能并展现出强大的泛化能力。

英文摘要

Building a generalized affordance grounding model to identify actionable regions on objects is vital for real-world applications. Existing methods to train the model can be divided into weakly and fully supervised ways. However, the former method requires a complex training framework design and can not infer new actions without an auxiliary prior. While the latter often struggle with limited annotated data and components trained from scratch despite being simpler. This study focuses on fully supervised affordance grounding and overcomes its limitations by proposing AffordanceSAM, which extends SAM's generalization capacity in segmentation to affordance grounding. Specifically, we design an affordance-adaption module and curate a coarse-to-fine annotated dataset called C2F-Aff to thoroughly transfer SAM's robust performance to affordance in a three-stage training manner. Experimental results confirm that AffordanceSAM achieves state-of-the-art (SOTA) performance on the AGD20K benchmark and exhibits strong generalized capacity.

URL PDF HTML 收藏
2607.07855 2026-07-10 cs.LG 新提交

NFTR: From Provable Mode-Averaging to Geodesic Subgoal Selection in Offline Goal-Conditioned RL

NFTR:从可证明的模式平均到离线目标条件强化学习中的测地线子目标选择

Erdemt Bao, Xing Lei, Jun Chen

机构 * Huazhong University of Science and Technology(华中科技大学) Xi’an Jiaotong University(西安交通大学) University of Electronic Science and Technology of China(电子科技大学)

AI总结 针对分层隐式Q学习在离线目标条件强化学习中存在的问题,提出NFTR方法,用条件归一化流取代高斯策略,结合基于架构三角不等式的三角松弛分数及RWDR目标,可避免高斯坍塌且在随机动力学下保持稳定。

详情
AI中文摘要

分层隐式Q学习(HIQL)是一种离线目标条件强化学习方法,仅通过价值函数优势来选择子目标。该规则有两种耦合的失败模式。乐观偏差将幸运的随机结果视为熟练的选择,而模式坍塌将多模态子目标分布减少到单个高斯均值,该均值通常落在无法到达的区域。我们提出了NFTR(具有三角松弛重加权的归一化流子目标策略)。条件归一化流取代了高斯策略,并且一个闭式模式平均结果将归一化流识别为基于AWR的子目标选择的最小生成类。基于架构三角不等式且不依赖距离准确性构建的三角松弛分数,通过乘法校正AWR权重,以降低迂回成本超过平均可达性的子目标的权重。三角松弛在确定性MDP的测地线上消失,并且在随机动力学下仍然是可组合性违反的保守上界。RWDR目标保留了AWR的总体水平单调改进,并允许进行三项次优分解。这两个要素共同产生了可证明地避免上述高斯坍塌且在随机动力学下保持稳定的子目标选择。

英文摘要

Hierarchical Implicit Q-Learning (HIQL), an offline goal-conditioned RL method, selects subgoals by value-function advantages alone. This rule has two coupled failure modes. Optimistic bias treats lucky stochastic outcomes as skillful choices, and mode collapse reduces a multi-modal subgoal distribution to a single Gaussian mean that often falls in unreachable regions. We propose NFTR (Normalizing Flows subgoal policies with Triangle-slack Reweighting). A conditional Normalizing Flow replaces the Gaussian policy, and a closed-form mode-averaging result identifies NFs as the minimal generative class for AWR-based subgoal selection. A triangle slack score, built on the architectural triangle inequality without relying on distance accuracy, multiplicatively corrects the AWR weight to downweight subgoals whose detour cost exceeds average reachability. Triangle-slack vanishes on geodesics in deterministic MDPs and remains a conservative upper bound on composability violation under stochastic dynamics. The RWDR objective preserves AWR's population-level monotonic improvement and admits a three-term suboptimality decomposition. Together, these two ingredients yield subgoal selection that provably avoids the Gaussian collapse described above and remains stable under stochastic dynamics. GitHub page: https://github.com/erdemtbao/NFTR

URL PDF HTML 收藏
2502.20805 2026-07-10 cs.RO cs.CV 版本更新

FunHOI: Annotation-Free 3D Hand-Object Interaction Generation via Functional Text Guidance

FunHOI:通过功能文本引导实现无注释的3D手-物体交互生成

Yongqi Tian, Xueyu Sun, Haoyuan He, jianlei Wang, Caigui Jiang

机构 * State Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) Xi’an Jiaotong University(西安交通大学)

AI总结 研究针对手-物体交互中功能抓取语义捕捉难的问题,提出两阶段框架FGS-Net,通过文本引导3D模型生成器FGG和姿态优化策略FGR,无需额外3D注释数据即可实现精确高质量的3D手-物体交互生成。

详情
AI中文摘要

手-物体交互(HOI)是人与环境的基本联系,但其灵巧复杂的姿态给手势控制带来重大挑战。尽管人工智能和机器人技术取得了显著进展,但捕捉功能抓取任务的语义仍是巨大挑战。此前工作虽能生成稳定正确的3D抓取,但因未考虑抓取语义,离功能抓取仍有差距。为应对这一挑战,我们提出创新的两阶段框架功能抓取合成网络(FGS-Net),由文本引导的3D模型生成器功能抓取生成器(FGG)和姿态优化策略功能抓取精炼器(FGR)组成,能基于文本输入生成3D HOI。大量实验表明该方法无需额外3D注释数据就能实现精确高质量的HOI生成。

英文摘要

Hand-object interaction(HOI) is the fundamental link between human and environment, yet its dexterous and complex pose significantly challenges for gesture control. Despite significant advances in AI and robotics, enabling machines to understand and simulate hand-object interactions, capturing the semantics of functional grasping tasks remains a considerable challenge. While previous work can generate stable and correct 3D grasps, they are still far from achieving functional grasps due to unconsidered grasp semantics. To address this challenge, we propose an innovative two-stage framework, Functional Grasp Synthesis Net (FGS-Net), for generating 3D HOI driven by functional text. This framework consists of a text-guided 3D model generator, Functional Grasp Generator (FGG), and a pose optimization strategy, Functional Grasp Refiner (FGR). FGG generates 3D models of hands and objects based on text input, while FGR fine-tunes the poses using Object Pose Approximator and energy functions to ensure the relative position between the hand and object aligns with human intent and remains physically plausible. Extensive experiments demonstrate that our approach achieves precise and high-quality HOI generation without requiring additional 3D annotation data.

URL PDF HTML 收藏
2607.06871 2026-07-09 cs.CV 新提交

Geometric Collapse: When Vision Models Fail to Verify Physical Causality

几何崩溃:视觉模型何时无法验证物理因果关系

Wentao Zhang, Jinhu Qi, Weiqiang Jin, Yifei Zhang, Chan-Tong Lam, Irwin King

机构 * Faculty of Applied Sciences, Macao Polytechnic University(澳门理工学院应用科学学院) The Chinese University of Hong Kong(香港中文大学) AgentecFusion Limited(AgentecFusion有限公司) School of Information and Communications Engineering, Xi’an Jiaotong University(西安交通大学信息与通信工程学院) School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院)

AI总结 研究视觉模型在验证物理因果关系时的问题,提出“Scrambled Edges”方法,通过特定控制分离无支撑边缘证据影响,实验表明该方法使预测偏差增大,几何崩溃全局传播,为改进模型物理合理性提供依据。

Comments ICML 2026

详情
AI中文摘要

大规模自监督学习的最新进展改进了密集几何预测,但尚不清楚这种扩展是否能在推理时进行物理合理性检查。我们提出了“Scrambled Edges”,这是一种可控的反事实方法,在违反表面连续性、光照一致性和遮挡顺序的同时注入突出的边缘状线索。通过能量匹配和结构匹配的控制,我们从高频能量和边缘稀疏性中分离出无支撑边缘证据的影响。在NYU Depth v2和KITTI上的CNN/ViT/SSL深度预测器中,“Scrambled Edges”导致的与干净预测的偏差比能量匹配噪声大3.2倍;额外的扩散和流匹配深度估计器显示偏差减弱但仍很显著。由此产生的几何崩溃会全局传播:即使知道损坏区域的准确信息,输出级修复也只能恢复47%,掩码外仍有大量误差。这些发现提供了可控的行为证据,表明当前的密集预测器缺乏可靠机制来隔离物理上无支撑的边缘线索,这促使进行明确的合理性评分和选择性线索整合。

英文摘要

Recent progress in large-scale self-supervised learning has improved dense geometric prediction, but it remains unclear whether such scaling yields inference-time physical plausibility checks. We propose Scrambled Edges, a controlled counterfactual that injects salient edge-like cues while violating surface continuity, illumination coherence, and occlusion ordering. With energy-matched and structure-matched controls, we isolate the effect of unsupported edge evidence from high-frequency energy and edge sparsity. Across CNN/ViT/SSL depth predictors on NYU Depth v2 and KITTI, Scrambled Edges induce up to 3.2x larger deviation from clean predictions than energy-matched noise; additional diffusion and flow-matching depth estimators show attenuated but still significant collapse. The resulting Geometric Collapse propagates globally: even with oracle knowledge of the corrupted region, output-level repair recovers only 47%, with substantial error outside the mask. These findings provide controlled behavioral evidence that current dense predictors lack reliable mechanisms to quarantine physically unsupported edge cues, motivating explicit plausibility scoring and selective cue integration.

URL PDF HTML 收藏
2512.05693 2026-07-09 cs.RO cs.AI 版本更新

HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies

HiMoE-VLA:用于通用视觉-语言-动作策略的分层专家混合模型

Zhiying Du, Bei Liu, Yaobo Liang, Yichao Shen, Haidong Cao, Xiangyu Zheng, Zhiyuan Feng, Zuxuan Wu, Jiaolong Yang, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Microsoft Research Asia(微软亚洲研究院) Xi’an Jiaotong University(西安交通大学) Tsinghua University(清华大学)

AI总结 研究通用视觉-语言-动作策略训练中异构性引发的负迁移问题,提出HiMoE-VLA框架,通过分层专家混合动作模块及两个辅助目标来解决,在多项任务上取得较好结果,还能将负迁移转化为正迁移。

详情
AI中文摘要

通用视觉-语言-动作(VLA)策略通常在跨越不同实体、动作空间和观察配置的机器人演示的异构混合上进行训练。使用共享密集动作模块对这种异构性进行建模可能会引发负迁移,特别是当动作空间或视觉观察在不同数据源之间存在差异时。我们使用HiMoE-VLA来解决这个问题,这是一个围绕分层专家混合(HiMoE)动作模块构建的VLA框架。HiMoE在输入/输出边界使用动作空间专家混合层,为不同动作空间专门化计算;在相邻层使用异构平衡专家混合层,为观察、场景和实体中的剩余差异提供平衡能力;在中间使用密集Transformer块来整合共享表示。两个辅助目标进一步指导这个层次结构:用于边界专门化的对比动作空间正则化目标和用于稳定专家利用的负载平衡目标。HiMoE-VLA在CALVIN上达到3.98,在LIBERO上达到98.0%,在真实xArm7和ALOHA任务上平均成功率分别为75.0%和63.7%;在受控异构协同训练下,它将在强基线中观察到的负迁移转化为正迁移。代码和模型可在指定网址公开获取。

英文摘要

Generalist vision--language--action (VLA) policies are typically trained on heterogeneous mixtures of robot demonstrations spanning diverse embodiments, action spaces, and observation configurations. Modeling such heterogeneity with a shared dense action module can induce negative transfer, particularly when action spaces or visual observations differ across data sources. We address this issue with HiMoE-VLA, a VLA framework built around a Hierarchical Mixture-of-Experts (HiMoE) action module. HiMoE uses Action-Space MoE layers at the input/output boundaries to specialize computation for distinct action spaces, Heterogeneity-Balancing MoE layers in neighboring layers to provide balanced capacity for residual variation in observations, scenes, and embodiments, and dense Transformer blocks in the middle to integrate shared representations. Two auxiliary objectives further guide this hierarchy: a contrastive Action-Space Regularization objective for boundary specialization and a load-balancing objective for stable expert utilization. HiMoE-VLA reaches 3.98 on CALVIN, 98.0\% on LIBERO, and 75.0\% and 63.7\% average success on real xArm7 and ALOHA tasks; under controlled heterogeneous co-training, it turns the negative transfer observed in strong baselines into positive transfer. The code and models are publicly available at https://github.com/ZhiyingDu/HiMoE-VLA.

URL PDF HTML 收藏
2607.06151 2026-07-08 cs.LG math.PR 新提交

Leveraging Extragradient for Effective Sharpness-Aware Minimization in Deep Learning

利用外梯度实现深度学习中有效的锐度感知最小化

Yao Fu, Chunxia Zhang, Junmin Liu, Yihang Jin, Haishan Ye, Yuanao Yang

机构 * SGIT AI Lab, State Grid Corporation of China(国网SGIT人工智能实验室) School of Management, Xi’an Jiaotong University(西安交通大学管理学院) School of Software, Xi’an Jiaotong University(西安交通大学软件学院)

AI总结 研究针对深度学习泛化难题,基于SAM提出EISAM优化器,通过两步更新过程提升泛化性能,对扰动半径敏感度降低。实验表明其在测试准确率和训练效率上优于多种优化器,理论分析证实其收紧泛化边界,还提供实用调优指导。

详情
AI中文摘要

泛化仍是深度学习中的关键挑战,传统优化器如随机梯度下降(SGD)常收敛到尖锐最小值,导致过拟合和对未见数据性能下降。基于锐度感知最小化(SAM),为寻求与改进泛化相关的平坦最小值,我们提出外梯度启发的锐度感知最小化(EISAM),一种通过外梯度技术增强泛化的新型优化器。EISAM采用两步更新过程:预测步骤研究损失景观几何,扰动步骤用基础优化器细化更新。此方法比SAM有更好的泛化性能,对扰动半径敏感度降低,增强了鲁棒性并简化了不同设置下的调优。在基准数据集上的大量实验表明,EISAM在各种架构的测试准确率和训练效率上始终优于SGD、自适应矩估计(Adam)和SAM。理论分析进一步证实,EISAM通过将参数导向曲率更小的更平坦最小值来收紧泛化边界。伴随全面的超参数分析,EISAM提供了实用的调优指导,成为推进深度学习理论和实践的强大、可扩展且广泛适用的优化解决方案。

英文摘要

Generalization remains a pivotal challenge in deep learning, where traditional optimizers like Stochastic Gradient Descent (SGD) often converge to sharp minima, leading to overfitting and reduced performance on unseen data. Building on Sharpness-Aware Minimization (SAM), for seeking flat minima associated with improved generalization, we propose the Extragradient-Inspired Sharpness-Aware Minimization (EISAM), a novel optimizer that enhances generalization via the extragradient technique. EISAM uses a two-step update process: a prediction step investigating the geometry of the loss landscape and a perturbation step that refines updates with a base optimizer. This approach achieves better generalization performance than SAM. Crucially, EISAM reduces sensitivity to the perturbation radius, enhancing robustness, and simplifying the tuning across diverse settings. Extensive experiments on benchmark datasets demonstrate that EISAM consistently outperforms SGD, Adaptive Moment Estimation (Adam), and SAM in test accuracy and training efficiency across various architectures. Theoretical analysis further confirms that EISAM tightens the generalization bound by steering parameters toward flatter minima with reduced curvature. Accompanied by a thorough hyperparameter analysis, EISAM offers practical tuning guidance, establishing it as a robust, scalable, and broadly applicable optimization solution that advances both the theory and practice in deep learning.

URL PDF HTML 收藏
2604.08435 2026-07-08 cs.CV cs.AI 版本更新

HST-HGN: Heterogeneous Spatial-Temporal Hypergraph Networks with Bidirectional State Space Models for Global Fatigue Assessment

HST-HGN:基于双向状态空间模型的异构空间-时间超图网络用于全局疲劳评估

Changdao Chen, Qinqiuhong Ye, Hao Chen, Jinyu Wang

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)

AI总结 本文提出HST-HGN,通过双向状态空间模型和异构空间-时间超图网络,有效建模长程时序依赖,实现对驾驶员疲劳的高效评估,兼顾判别力与计算效率,适用于实时车载边缘部署。

Comments 10 pages

详情
AI中文摘要

本文提出HST-HGN,通过双向状态空间模型和异构空间-时间超图网络,有效建模长程时序依赖,实现对驾驶员疲劳的高效评估,兼顾判别力与计算效率,适用于实时车载边缘部署。

英文摘要

It remains challenging to assess driver fatigue from untrimmed videos under constrained computational budgets, due to the difficulty of modeling long-range temporal dependencies in subtle facial expressions. Some existing approaches rely on computationally heavy architectures, whereas others employ traditional lightweight pairwise graph networks, despite their limited capacity to model high-order synergies and global temporal context. Therefore, we propose HST-HGN, a novel Heterogeneous Spatial-Temporal Hypergraph Network driven by Bidirectional State Space Models. Spatially, we introduce a hierarchical hypergraph network to fuse pose-disentangled geometric topologies with multi-modal texture patches dynamically. This formulation encapsulates high-order synergistic facial deformations, effectively overcoming the limitations of conventional methods. In temporal terms, a Bi-Mamba module with linear complexity is applied to perform bidirectional sequence modeling. This explicit temporal-evolution filtering enables the network to distinguish highly ambiguous transient actions, such as yawning versus speaking, while encompassing their complete physiological lifecycles. Extensive evaluations across diverse fatigue benchmarks demonstrate that HST-HGN achieves state-of-the-art performance. In particular, our method strikes a balance between discriminative power and computational efficiency, making it well-suited for real-time in-cabin edge deployment.

URL PDF HTML 收藏
2112.02353 2026-07-08 cs.CV cs.LG 版本更新

Label Hierarchy Transition: Delving into Class Hierarchies to Enhance Deep Classifiers

标签层次转换:深入研究类层次结构以增强深度分类器

Renzhen Wang, De cai, Kaiwen Xiao, Xixi Jia, Xiao Han, Deyu Meng

机构 * School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(数学与统计学学院和教育部智能网络与网络安全重点实验室,西安交通大学) Tencent AI Lab(腾讯AI实验室) ByteDance(字节跳动) SINGPATH AI Lab, SingPath Medical Technology Pte. Ltd.(SingPath AI实验室,SingPath医疗科技私人有限公司) School of Mathematics and Statistics, Xidian University(数学与统计学学院,西安电子科技大学) College of Biomedical Engineering, Sichuan University(生物医学工程学院,四川大学) School of Mathematics and Statistics, Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(数学与统计学学院和教育部智能网络与网络安全重点实验室,西安交通大学) Macao Institute of Systems Engineering, Macau University of Science and Technology(澳门系统工程研究院,澳门科技大学)

AI总结 研究针对层次分类中现有方法未充分利用类别相关性的问题,提出基于深度学习的LHT统一概率框架,由转换网络和混淆损失构成,实验证明其优于现有方法,还扩展到皮肤病变诊断任务展现潜力。

详情
AI中文摘要

层次分类旨在将对象分类到层次化的类别结构中,如鸟类可按目、科、种的三级层次分类。现有方法常将其解耦为一系列多类分类任务,但这种多任务学习策略未能充分利用层次结构中不同级别各类别间的相关性。本文提出基于深度学习的统一概率框架标签层次转换(LHT)来应对层次分类挑战。LHT框架由转换网络和混淆损失组成,转换网络专注于显式学习标签层次转换矩阵,可有效编码类层次结构中的潜在相关性,混淆损失促使分类网络在训练中学习不同标签层次间的相关性。该框架只需少量修改就能适应任何现有深度网络。通过一系列公共基准数据集进行层次分类问题实验,结果表明该方法优于当前最先进方法。此外,还将LHT框架扩展到皮肤病变诊断任务,验证了其在计算机辅助诊断中的巨大潜力。方法代码可在指定链接获取。

英文摘要

Hierarchical classification aims to sort the object into a hierarchical structure of categories. For example, a bird can be categorized according to a three-level hierarchy of order, family, and species. Existing methods commonly address hierarchical classification by decoupling it into a series of multi-class classification tasks. However, such a multi-task learning strategy fails to fully exploit the correlation among various categories across different levels of the hierarchy. In this paper, we propose Label Hierarchy Transition (LHT), a unified probabilistic framework based on deep learning, to address the challenges of hierarchical classification. The LHT framework consists of a transition network and a confusion loss. The transition network focuses on explicitly learning the label hierarchy transition matrices, which has the potential to effectively encode the underlying correlations embedded within class hierarchies. The confusion loss encourages the classification network to learn correlations across different label hierarchies during training. The proposed framework can be readily adapted to any existing deep network with only minor modifications. We experiment with a series of public benchmark datasets for hierarchical classification problems, and the results demonstrate the superiority of our approach beyond current state-of-the-art methods. Furthermore, we extend our proposed LHT framework to the skin lesion diagnosis task and validate its great potential in computer-aided diagnosis. The code of our method is available at \href{https://github.com/renzhenwang/label-hierarchy-transition}{https://github.com/renzhenwang/label-hierarchy-transition}.

URL PDF HTML 收藏
2607.02845 2026-07-07 cs.RO cs.AI cs.CV 新提交

Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation

受差分放大器启发的AmpAttention用于多视图机器人操作

Jin Yang, Ping Wei, Nanning Zheng

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, and Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人机混合增强智能技术国家级重点实验室、人工智能与机器人研究所)

AI总结 针对机器人视图图像问题,提出受模拟电路差分放大器启发的AmpAttention机制,抑制注意力噪声。基于此引入RVAF模型,提升训练效率与任务性能,还扩展为RVAF++,在高精度任务上成果显著。

Comments Accepted by IROS2026

详情
AI中文摘要

具有注意力机制的多视图机器人操作方法在训练效率和任务性能方面取得了显著进展。然而,机器人视图图像中固有的冗余、遮挡和视点依赖性常常导致严重的注意力漂移。为了应对这一挑战,我们提出了AmpAttention,一种受模拟电路中的差分放大器启发的新型注意力机制。它旨在抑制注意力噪声并捕获高信噪比信号以实现更可靠的感知。基于此,我们引入了RVAF模型,该模型集成了任务引导的视图内和视图间AmpAttention。与先前的最先进方法相比,RVAF在18个RLBench任务(249个变体)中实现了最佳平均成功率,同时将训练时间减少了33.3%。RVAF在现实世界的高精度任务中也显示出强大的潜力,例如它能够拿起飞镖并准确地将其插入红色靶心。此外,我们通过合并SAM2图像编码器将RVAF扩展到RVAF++。RVAF++在高精度任务上取得了显著进展,在“插入钉子”任务上实现了91%的成功率。更多定性结果可在匿名项目网站this https URL上获得。

英文摘要

Multi-view robotic manipulation methods with the attention mechanism have recently achieved significant progress in both training efficiency and task performance. However, the inherent redundancy, occlusion, and viewpoint dependency in robotic view images often lead to severe attention drift. To address this challenge, we propose AmpAttention, a novel attention mechanism inspired by differential amplifiers in analog circuits. It aims to suppress attention noise and capture high signal-to-noise ratio signals for more reliable perception. Based on this, we introduce the RVAF model, which integrates task-guided intra-view and inter-view AmpAttention. Compared to previous state-of-the-art methods, RVAF achieves the optimal average success rate across 18 RLBench tasks (249 variations) while reducing training time by 33.3\%. RVAF also demonstrates strong potential in real-world high-precision tasks, exemplified by its ability to pick up a dart and accurately insert it into the red bullseye. Furthermore, we extend RVAF to RVAF++ by incorporating the SAM2 image encoder. RVAF++ achieves substantial gains on high-precision tasks, achieving a 91\% success rate on the `insert peg' task. More qualitative results are provided at the anonymous project website https://anonymous.4open.science/w/RVAF-Anonymization.

URL PDF HTML 收藏
2602.19543 2026-07-07 cs.CL cs.IR 版本更新

Hyper-KGGen: A Skill-Driven Knowledge Extractor for High-Quality Knowledge Hypergraph Generation

Hyper-KGGen:一种用于高质量知识超图生成的技能驱动知识提取器

Rizhuo Huang, Yifan Feng, Rundong Xue, Shihui Ying, Jun-Hai Yong, Chuan Shi, Shaoyi Du, Yue Gao

机构 * Xi’an Jiaotong University(西安交通大学) State Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) Tsinghua University(清华大学) Shanghai University(上海大学) Beijing University of Posts and Telecommunications(北京邮电大学)

AI总结 针对高质量知识超图构建难题,提出Hyper-KGGen框架,用粗到细机制分解文档,含自适应技能获取模块,经反馈回路提炼领域专长,还给出标注基准,实验验证其性能优于基线。

详情
AI中文摘要

知识超图通过封装复杂n元原子事实超越传统二元知识图,为语义表示提供更全面范式。但构建高质量超图因场景差距仍具挑战。我们提出Hyper-KGGen,将提取重新表述为动态技能演化过程,采用粗到细机制并结合自适应技能获取模块,经基于稳定性的反馈回路提炼领域专长。还给出HyperDocRED基准。实验表明Hyper-KGGen显著优于基线。

英文摘要

Knowledge hypergraphs surpass traditional binary knowledge graphs by encapsulating complex n-ary atomic facts, providing a more comprehensive paradigm for semantic representation. However, constructing high-quality hypergraphs remains challenging due to the scenario gap: generic extractors struggle to generalize across diverse domains with specific jargon, while existing methods often fail to balance structural skeletons with fine-grained details. To bridge this gap, we propose Hyper-KGGen, a skill-driven framework that reformulates extraction as a dynamic skill-evolving process. First, Hyper-KGGen employs a coarse-to-fine mechanism to systematically decompose documents, ensuring full-dimensional coverage from binary links to complex hyperedges. Crucially, it incorporates an adaptive skill acquisition module that actively distills domain expertise into a Global Skill Library. This is achieved via a stability-based feedback loop, where extraction stability serves as a relative reward signal to induce high-quality skills from unstable traces and missed predictions. Additionally, we present HyperDocRED, a rigorously annotated benchmark for document-level knowledge hypergraph extraction. Experiments demonstrate that Hyper-KGGen significantly outperforms strong baselines, validating that evolved skills provide substantially richer guidance than static few-shot examples in multi-scenario settings.

URL PDF HTML 收藏