arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Korea Advanced Institute of Science and Technology(韩国科学技术院)

至 收录 1292
2607.18119 2026-07-21 stat.ML cs.LG 新提交

COVAriance-Induced Fairness Gap Penalty for Subgroup-Fair Clustering

用于子群公平聚类的协方差诱导公平差距惩罚

Kyungseon Lee, Hankyo Jeong, Kunwoong Kim, Kwanho Lee, Yongdai Kim

机构 * Department of Statistics, Seoul National University(首尔国立大学统计系) KAIST AI(韩国科学技术院人工智能研究所)

AI总结 研究在多敏感属性定义多个子群时的公平聚类问题,提出基于协方差诱导公平差距惩罚的方法,推导代理并连续松弛,扩展框架捕捉子群 - 边际公平差距,实验证明算法在成本 - 公平权衡上有竞争力且提高计算效率。

详情
AI中文摘要

公平聚类旨在使聚类分配与敏感属性无关,但当多个敏感属性共同定义多个子群时,这一目标颇具挑战。直接扩展现有公平聚类算法计算成本高或数值不稳定,尤其是子群数量呈指数增长且部分子群实例很少时。为应对这些挑战,我们定义了聚类的子群公平差距并推导了基于协方差的代理,它与该差距完全匹配。接着引入代理的连续松弛,实现基于梯度的高效优化并得出算法COVA - FC。我们还表明子群公平并不意味着边际公平,并扩展框架以捕捉子群 - 边际公平差距。基准数据集实验表明,COVA - FC在成本 - 公平权衡方面具有竞争力,且在子群和高阶边际设置中提高了计算效率。

英文摘要

Fair clustering aims to make cluster assignments independent of sensitive attributes, but this goal becomes challenging when multiple sensitive attributes jointly define many subgroups. In such settings, directly extending existing fair clustering algorithms is computationally expensive or numerically unstable, especially when the number of subgroups grows exponentially and some subgroups contain only a few instances. To address these challenges, we define a subgroup-fairness gap for clustering and derive a covariance-based surrogate that exactly matches this gap. We then introduce a continuous relaxation of the surrogate, enabling efficient gradient-based optimization and yielding our proposed algorithm, COVA-FC. We also show that subgroup fairness alone does not imply marginal fairness, and extend our framework to capture a subgroup-marginal-fairness gap. Experiments on benchmark datasets show that COVA-FC achieves competitive cost-fairness trade-offs and improves computational efficiency over existing baselines in both subgroup and higher-order marginal settings.

URL PDF HTML 收藏
2607.17272 2026-07-21 cs.LG cs.SI 新提交

Node4All: Learning Node Representation Beyond Datasets

Node4All:超越数据集学习节点表示

Dooho Lee, Jaemin Yoo

机构 * KAIST(韩国科学技术院) Seoul National University(首尔国立大学)

AI总结 研究针对多数节点表示学习方法依赖特定数据集训练的问题,提出Node4All,基于通道图变换器和自监督学习构建,能跨任意图数据集泛化,在节点分类任务中表现出色,实现可复用性且实践效果好。

Comments Accepted to KDD 2026

详情
AI中文摘要

节点表示学习发展迅速,但多数现有方法依赖于每个数据集的训练和超参数调整。这种特定于数据集的优化源于设计可跨不同图数据集泛化的可复用图模型的困难。本文介绍了Node4All,一种无需任何特定于数据集的优化即可应用于任意图数据集的节点表示学习器。它基于两个互补思想构建。在架构层面,引入通道图变换器(CGT),能以单一固定参数化处理任意图数据集。在学习层面,提出基于一系列合成图的自监督学习。通过广泛评估,Node4All在25个基准测试的节点分类任务中与21个基线方法竞争,排名第5,还支持一次性和上下文学习,优于近期图基础模型。这些结果表明Node4All不仅实现了跨任意图数据集的可复用性,在实践中也是有效的解决方案。

英文摘要

Node representation learning has advanced rapidly, yet most existing methods rely on per-dataset training and hyperparameter tuning. This dataset-specific optimization comes from the difficulty of designing reusable graph models that generalize across diverse graph datasets. In this work, we introduce Node4All, a node representation learner applicable to arbitrary graph datasets without any dataset-specific optimization. Node4All is built on two complementary ideas. At the architectural level, we introduce the Channel Graph Transformer (CGT), which enables a single fixed parameterization to process arbitrary graph datasets. At the learning level, we propose a self-supervised learning based on a series of synthetic graphs. Together, these components enable generalization beyond individual datasets, which is infeasible with existing architectures and learning frameworks. We extensively evaluate Node4All on node classification across 25 benchmarks against 21 baselines, covering both supervised and self-supervised methods. Despite all baselines being trained and optimized for each dataset, a single Node4All, applied uniformly across the datasets, achieves a competitive ranking of 5th among 21 baselines. Moreover, Node4All supports one-shot and in-context learning with an appropriate predictor and outperforms recent graph foundation models (GFMs) in these settings. These results demonstrate that Node4All not only achieves reusability across arbitrary graph datasets, but also remains an effective solution in practice. Code and model checkpoints are available in https://github.com/dooho00/node4all.

URL PDF HTML 收藏
2607.16926 2026-07-21 cs.CV 新提交

Splat-based 3D Scene Reconstruction with Extreme Motion-blur

基于体素的极端运动模糊3D场景重建

Hyeonjoong Jang, Dongyoung Choi, Donggun Kim, Woohyun Kang, Min H. Kim

机构 * KAIST(韩国科学技术院) HYPERGRAM(超图)

AI总结 研究针对低光照下RGB-D输入的极端运动模糊问题,提出结合高斯体素框架的相机姿态估计与图像去模糊方法,通过对齐帧、调整高斯位置等步骤,提升3D重建质量,优于现有方法,有广泛应用意义。

Journal ref Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV) 2025

详情
AI中文摘要

我们提出一种基于体素的3D场景重建方法,用于处理RGB-D输入中的极端运动模糊,这在低光照环境中是一个常见挑战。在昏暗照明下,RGB帧常因曝光时间长而出现严重运动模糊,导致传统相机姿态估计方法失败,影响3D重建质量。尽管近期技术有成果,但在快速运动或光照不佳时相机轨迹估计困难。我们引入结合相机姿态估计和图像去模糊的方法,利用高斯体素和深度输入增强场景表示。先通过光流和ICP对齐连续RGB-D帧,再调整高斯位置优化深度对齐来细化相机姿态和3D几何。通过比较输入与一系列清晰渲染帧去模糊图像。实验表明该方法优于现有方法,对机器人、自主导航和增强现实中的3D映射应用有广泛意义,代码和数据集公开。

英文摘要

We propose a splat-based 3D scene reconstruction method from RGB-D input that effectively handles extreme motion blur, a frequent challenge in low-light environments. Under dim illumination, RGB frames often suffer from severe motion blur due to extended exposure times, causing traditional camera pose estimation methods, such as COLMAP, to fail. This results in inaccurate camera pose and blurry color input, compromising the quality of 3D reconstructions. Although recent 3D reconstruction techniques like Neural Radiance Fields and Gaussian Splatting have demonstrated impressive results, they rely on accurate camera trajectory estimation, which becomes challenging under fast motion or poor lighting conditions. Furthermore, rapid camera movement and the limited field of view of depth sensors reduce point cloud overlap, limiting the effectiveness of pose estimation with the ICP algorithm. To address these issues, we introduce a method that combines camera pose estimation and image deblurring using a Gaussian Splatting framework, leveraging both 3D Gaussian splats and depth inputs for enhanced scene representation. Our method first aligns consecutive RGB-D frames through optical flow and ICP, then refines camera poses and 3D geometry by adjusting Gaussian positions for optimal depth alignment. To handle motion blur, we model camera movement during exposure and deblur images by comparing the input with a series of sharp, rendered frames. Experiments on a new RGB-D dataset with extreme motion blur show that our method outperforms existing approaches, enabling high-quality reconstructions even in challenging conditions. This approach has broad implications for 3D mapping applications in robotics, autonomous navigation, and augmented reality. Both code and dataset are publicly available on https://github.com/KAIST-VCLAB/gs-extreme-motion-blur.

URL PDF HTML 收藏
2607.03788 2026-07-21 cs.LG 版本更新

Tensor-Train Joint Modeling for Few-Step Discrete Diffusion

用于少步离散扩散的张量训练联合建模

Byoungkwon Kim, Minhyuk Sung

机构 * KAIST(韩国科学技术院)

AI总结 研究离散扩散少步生成潜力受限问题,用张量分解进行联合分布建模,支持多种分解方式,提出迭代边际推理程序,通过微调预训练模型大幅提升少步生成效果。

详情
AI中文摘要

离散扩散有望比自回归模型更快生成顺序离散数据,但由于结构限制其少步生成潜力未被充分挖掘。本文提出首个通过张量分解进行离散扩散显式联合分布建模的框架,支持多种分解方式,识别出TTD对附近令牌依赖的结构偏差,提出迭代边际推理程序,通过微调预训练模型提升少步生成效果。

英文摘要

Discrete diffusion promises orders-of-magnitude faster generation than autoregressive (AR) models for sequential discrete data, yet its full potential of few-step generation has remained out of reach due to a fundamental structural limitation. The conditional-independence assumption underlying current discrete diffusion models introduces a systematic parallelization bias that compounds with the number of tokens unmasked per step, becoming severe in the few-step regime that fast generation requires. We address this with the first framework for explicit joint distribution modeling in discrete diffusion via tensor decomposition, which represents the conditional clean distribution as a low-rank tensor with controllable expressivity. The framework supports both Canonical Polyadic (CPD) and Tensor-Train (TTD) decompositions, and we identify a structural bias of TTD toward dependencies between nearby tokens, formalized through Oseledets' theorem relating TT-rank to unfolding-matrix rank, which is well-suited to sequential data such as natural language and line notations for molecular data. To enable efficient generation, we present an iterative marginal inference procedure with specialization for predetermined position schedules. Our framework integrates into pretrained MDMs through lightweight fine-tuning, yielding substantial improvements in few-step generation at a fraction of the cost of training from scratch. Code available at https://github.com/ssamt/tensor-train.

URL PDF HTML 收藏
2505.02722 2026-07-21 cs.AI cs.LG 版本更新

Enhancing LLMs' Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry

利用全国脓毒症登记处的真实世界数据增强大语言模型的临床推理能力

Junu Kim, Chaeeun Shim, Sungjin Park, Su Yeon Lee, Gee Young Suh, Chae-Man Lim, Seong Jin Choi, Song Mi Moon, Kyoung-Ho Song, Eu Suk Kim, Hong Bin Kim, Sejoong Kim, Chami Im, Dong-Wan Kang, Yong Soo Kim, Hee-Joon Bae, Sung Yoon Lim, Han-Gil Jeong, Edward Choi

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院) Microsoft(微软) Asan Medical Center, University of Ulsan College of Medicine(釜山大学医学院阿桑医疗中心) Samsung Medical Center, Sungkyunkwan University School of Medicine(成均馆大学医学院三星医疗中心) Seoul National University Bundang Hospital, Seoul National University College of Medicine(首尔国立大学医学院首尔国立大学医院)

AI总结 研究旨在增强大语言模型临床推理能力,利用全国脓毒症登记处数据构建问题,通过强化学习微调模型得到C-Reason,该模型在域内测试集表现出色,推理能力能推广到不同任务和疾病,为开发通用临床推理模型提供思路。

Comments Accepted at MLHC 2026

详情
AI中文摘要

尽管大语言模型(LLMs)在一般领域展现出令人印象深刻的推理能力,但在实际临床实践中的有效性仍有限。这可能是由于训练期间对真实世界临床数据接触不足,因隐私问题此类数据通常未被纳入。为解决此问题,我们提议利用真实世界临床数据增强LLMs的临床推理能力。我们从全国脓毒症登记处构建推理密集型问题,并使用强化学习在这些问题上对Phi-4进行微调,得到C-Reason。C-Reason在域内测试集上展现出强大的临床推理能力,通过定量指标和专家评估得以证明。此外,其增强的推理能力可推广到涉及不同任务和患者队列的脓毒症数据集、抗生素使用任务的开放式咨询以及其他疾病。未来研究应专注于用大规模、多疾病临床数据集训练LLMs,以开发更强大的通用临床推理模型。

英文摘要

Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in real-world clinical practice remains limited. This is likely due to their insufficient exposure to real-world clinical data during training, as such data is typically not included due to privacy concerns. To address this, we propose enhancing the clinical reasoning capabilities of LLMs by leveraging real-world clinical data. We constructed reasoning-intensive questions from a nationwide sepsis registry and fine-tuned Phi-4 on these questions using reinforcement learning, resulting in C-Reason. C-Reason exhibited strong clinical reasoning capabilities on the in-domain test set, as evidenced by both quantitative metrics and expert evaluations. Furthermore, its enhanced reasoning capabilities generalized to a sepsis dataset involving different tasks and patient cohorts, an open-ended consultations on antibiotics use task, and other diseases. Future research should focus on training LLMs with large-scale, multi-disease clinical datasets to develop more powerful, general-purpose clinical reasoning models.

URL PDF HTML 收藏
2607.16012 2026-07-20 cs.CV cs.AI cs.RO 新提交

DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction

DPNeXt:用于基于高效视觉Transformer的多任务密集预测的轻量级多尺度特征融合框架

Jehun Kang, Jungha Wang, Youngjun Hwang, David Hyunchul Shim

机构 * School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院电气工程学院)

AI总结 研究针对机器人感知系统多任务学习中现有解码策略瓶颈,提出DPNeXt框架,用双深度可分离倒置瓶颈及MTBG策略,在多任务密集预测上表现出色,相比DPT大幅减少参数并提升推理速度。

Comments 8 pages, 5 figures. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

详情
AI中文摘要

机器人感知系统中的多任务学习(MTL)通过整合语义分割和深度估计来支持对3D空间场景的全面理解。虽然视觉基础模型(VFM)越来越多地被用作强大的特征编码器,但现有的解码策略是一个关键瓶颈。为解决此问题,我们提出了DPNeXt,这是一种简化的多尺度特征融合解码器,是标准密集预测Transformer(DPT)的有效替代方案。DPNeXt使用双深度可分离倒置瓶颈,通过以融合为中心的解码和独立任务模块化来提高冻结VFM的利用率。为进一步减轻任务之间的负归纳转移,我们引入了多任务边界引导(MTBG)策略。与添加融合模块或门控的先前边界感知方法不同,MTBG应用对称的以边界为重点的监督来鼓励几何一致性,而无需额外的注释或推理成本。在Cityscapes上的实验表明,DPNeXt-S优于先前的最新(SOTA)MTL模型,而DPNeXt-B进一步提高了整体性能,并在比较方法中取得了最佳结果。在NYUv2上,DPNeXt-B在比较方法中也取得了最佳的语义分割和深度估计结果,同时所需的可训练参数比先前的大规模MTL模型少得多。与标准DPT相比,DPNeXt-S减少了78.6%的可训练参数,并在资源受限的笔记本电脑硬件上的比较模型中实现了最快的推理速度。源代码、模型检查点和演示视频将在这个https URL上提供。

英文摘要

Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation. While Vision Foundation Models (VFMs) are increasingly adopted as robust feature encoders, existing decoding strategies present a critical bottleneck. To address this, we propose DPNeXt, a streamlined multi-scale feature fusion decoder and efficient alternative to the standard Dense Prediction Transformer (DPT). DPNeXt uses dual depthwise separable inverted bottlenecks to improve frozen VFM utilization through fusion-centric decoding and independent task modularization. To further mitigate negative inductive transfer between tasks, we introduce the Multi-Task Boundary Guidance (MTBG) strategy. Unlike prior boundary-aware methods that add fusion modules or gating, MTBG applies symmetric boundary-focused supervision to encourage geometric consistency without extra annotation or inference cost. Experiments on Cityscapes show that DPNeXt-S outperforms prior state-of-the-art (SOTA) MTL models, while DPNeXt-B further improves the overall performance and achieves the best results among the compared methods. On NYUv2, DPNeXt-B also achieves the best semantic segmentation and depth estimation results among the compared methods while requiring substantially fewer trainable parameters than prior large-scale MTL models. Compared with the standard DPT, DPNeXt-S reduces trainable parameters by 78.6% and achieves the fastest inference speed among the compared models on resource-constrained laptop hardware. The source code, model checkpoints, and a demo video will be made available at https://github.com/kangjehun/DPNeXt.

URL PDF HTML 收藏
2607.15799 2026-07-20 cs.LG cs.AI 新提交

Knowledge-Assisted Multi-Graph Dependency Learning for Multivariate Time Series Anomaly Detection in Multi-Stage Industrial Processes

用于多阶段工业过程中多变量时间序列异常检测的知识辅助多图依赖学习

Jaeyeong Lee, Taeseong Yoon, Wonmo Koo, Heeyoung Kim

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

AI总结 针对多阶段工业过程中多变量时间序列异常检测问题,提出知识辅助多图框架并构建三个互补图,利用多图注意力网络建模,通过纳入过程知识显著提升异常检测性能。

详情
AI中文摘要

工业过程常常从多个阶段的多个传感器生成复杂且相互依赖的时间序列数据,变量和过程阶段之间形成复杂依赖关系。通过多变量时间序列异常检测(MTAD)对这些时间序列进行有效监测和及时异常检测,对于防止故障和确保自动化系统可靠性至关重要。图神经网络(GNN)通过利用数据驱动的图来建模变量间复杂依赖关系,推动了MTAD发展,但现有基于GNN的方法常忽略关键过程知识,即便考虑该知识,将其无缝融入现有模型也颇具挑战,导致性能欠佳。为解决此局限,我们提出一种知识辅助多图框架,用于多阶段工业过程中MTAD的传感器依赖建模,将过程知识明确纳入图学习以增强依赖建模并提升异常检测性能。我们的方法构建三个互补图:一个纯数据驱动图和两个通过整合过程知识导出的结构约束进行细化的图。为有效利用这些图进行异常检测,我们采用多图注意力网络,实现对复杂依赖关系更准确、稳健的表示。在两个真实世界的多阶段工业数据集上的综合实验表明,纳入过程知识可显著提升异常检测性能。

英文摘要

Industrial processes often generate complex, interdependent time-series data from multiple sensors across multiple stages, forming complex dependencies among variables and process stages. Effective monitoring and timely anomaly detection of these time series through multivariate time series anomaly detection (MTAD) is crucial for preventing failures and ensuring the reliability of automated systems. Graph neural networks (GNNs) have advanced MTAD by leveraging data-driven graphs to model complex dependencies among variables, effectively capturing relational structures within multivariate time series to enhance anomaly detection performance. However, existing GNN-based approaches often overlook critical process knowledge, and even when this knowledge is considered, seamlessly incorporating it into existing models remains inherently challenging, leading to suboptimal performance. To address this limitation, we propose a knowledge-assisted multi-graph framework for modeling sensor dependencies in multi-stage industrial processes for MTAD, which explicitly incorporates process knowledge into graph learning to enhance dependency modeling and improve anomaly detection performance. Our method constructs three complementary graphs: one purely data-driven and two refined by integrating structural constraints derived from process knowledge. To effectively leverage these graphs for anomaly detection, we employ a multi-graph attention network, enabling a more accurate and robust representation of complex dependencies. Comprehensive experiments on two real-world, multi-stage industrial datasets demonstrate that incorporating process knowledge substantially enhances anomaly detection performance.

URL PDF HTML 收藏
2607.15699 2026-07-20 cs.CV 新提交

GoStop: Reinforcement Learning for Adaptive Temporal Aggregation in Event-Based Feature Tracking

GoStop:基于强化学习的事件驱动特征跟踪中的自适应时间聚合

Youngho Kim, Hoonhee Cho, Jae-Young Kang, Kuk-Jin Yoon

机构 * KAIST(韩国科学技术院)

AI总结 研究基于事件相机的特征跟踪问题,提出用强化学习框架自适应控制事件累积过程,训练智能体依运动线索决策,引入新数据集评估,集成该框架到现有方法可提升性能,在动态运动下更具鲁棒性且平衡跟踪精度与效率。

Comments Accepted to ECCV 2026

详情
AI中文摘要

特征跟踪在理解场景运动中起着基础性作用,并支持各种下游任务。事件相机具有高时间分辨率和异步传感能力,适用于快速和非线性运动下的特征跟踪。然而,现有基于事件的特征跟踪方法依赖于基于手动调整的固定启发式规则进行事件累积,无法适应多样的运动动态。本文将事件累积建模为顺序决策问题,引入强化学习框架来自适应控制基于事件的在线特征跟踪的累积过程,并训练一个强化学习智能体根据运动线索决定是否继续累积事件或进行跟踪推理。此外,引入了具有动态运动分布的动态事件跟踪(DEFT)数据集来评估特征跟踪的鲁棒性。实验表明,将该即插即用框架集成到现有特征跟踪方法中始终优于基于启发式的方法,在动态运动下提高了鲁棒性,同时在跟踪准确性和效率之间取得了更好的平衡。

英文摘要

Feature tracking plays a fundamental role in understanding scene motion and supports various downstream tasks. Event cameras, with their high temporal resolution and asynchronous sensing, enable low-latency and motion-robust perception, making them well-suited for feature tracking under fast and non-linear motion. However, existing event-based feature tracking methods rely on fixed heuristic rules based on hand-tuning for event accumulation. Such strategies fail to adapt to diverse motion dynamics, leading to degraded performance under abrupt motion changes or low-motion scenarios. In this paper, we model event accumulation as a sequential decision-making problem and introduce reinforcement learning (RL) framework to adaptively control the accumulation process for online event-based feature tracking. Our approach trains a RL agent that decides whether to continue accumulating events or to perform tracking inference based on motion cues. The proposed adaptive temporal agent enables dynamic adaptation to varying motion patterns without relying on hand-crafted rules. Furthermore, we introduce a Dynamic Event-based Tracking (DEFT) dataset with dynamic motion distributions to evaluate the robustness of the feature tracking. Extensive experiments demonstrate that integrating our plug-and-play framework to existing feature tracking methods consistently outperforms heuristic-based approaches, improving robustness under dynamic motion while offering a better balance between tracking accuracy and efficiency. Our project codes and datasets are available at https://github.com/kmax2001/GoSTOP

URL PDF HTML 收藏
2607.16030 2026-07-20 quant-ph cs.AI cs.LG 新提交

Rethinking Quantum Continual Learning with Quantum Fisher Information

用量子费舍尔信息重新思考量子持续学习

Yu-Chao Hsu, Yu-Cheng Lin, Tai-Yue Li, Nan-Yow Chen, En-Jui Kuo

机构 * School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院电气工程学院) Cross College Elite Program, National Cheng Kung University(成功学院精英计划,成功大学) National Center for High-Performance Computing , National Institutes of Applied Research (NIAR)(高性能计算中心,应用研究国立机构) Department of Electrophysics, National Yang Ming Chiao Tung University(电子物理系,交通大学)

AI总结 研究量子持续学习中变分量子分类器的灾难性遗忘问题,提出基于量子费舍尔信息的QEWC方法,通过量化参数化量子态内在敏感性减轻遗忘,实验表明该方法能改善任务保留,在量子态几何研究中有重要意义。

详情
AI中文摘要

量子持续学习旨在训练量子模型处理序列任务而不丢失先前学到的知识。然而,变分量子分类器在非平稳任务分布下容易出现灾难性遗忘。我们提出了量子弹性权重巩固(QEWC),一种基于量子费舍尔信息(QFI)的正则化方法来减轻遗忘。与基于经典费舍尔信息(CFI)的传统弹性权重巩固不同,QEWC使用QFI量化参数化量子态的内在敏感性。我们在序列二分类任务训练的VQCs上评估QEWC,包括经典图像分类和量子相位分类任务。模拟表明无正则化的序列训练会导致严重遗忘,而基于CFI的EWC和基于QFI的QEWC都能改善对先前任务的保留。机理分析进一步表明两种方法施加不同的正则化几何:CFI选择性地作用于测量敏感方向,而QFI在参数空间上施加更密集的状态几何约束。在去极化噪声下,CFI值因测量统计退化而强烈抑制,而QFI保留噪声参数化量子态更稳定的敏感性结构。这些结果确立了QEWC作为一种通过量子态几何研究和减轻量子持续学习中遗忘的物理动机方法

英文摘要

Quantum continual learning aims to train quantum models on sequential tasks without losing previously learned knowledge. However, variational quantum classifiers (VQCs) are prone to catastrophic forgetting under nonstationary task distributions. We propose quantum elastic weight consolidation (QEWC), a quantum Fisher information (QFI)-informed regularization method for mitigating forgetting. Unlike conventional elastic weight consolidation based on classical Fisher information (CFI), which measures parameter importance through measurement-dependent output statistics, QEWC uses QFI to quantify the intrinsic sensitivity of the parameterized quantum state. This gives an information-geometric view in which important parameters are identified by the local response of the quantum state manifold. We evaluate QEWC on VQCs trained on sequential binary classification tasks, including classical image-classification and quantum phase-classification tasks. Simulations show that sequential training without regularization causes severe forgetting, while both CFI-based EWC and QFI-based QEWC improve retention of previous tasks. Mechanistic analyses further show that the two methods impose different regularization geometries: CFI acts selectively on measurement-sensitive directions, whereas QFI imposes a denser state-geometric constraint over parameter space. Under depolarizing noise, CFI values are strongly suppressed by degraded measurement statistics, while QFI preserves a more stable sensitivity structure of the noisy parameterized quantum state. These results establish QEWC as a physically motivated approach for studying and mitigating forgetting in quantum continual learning through quantum-state geometry.

URL PDF HTML 收藏
2607.14927 2026-07-17 cs.CV 新提交

TanGO: Training-Free 3D Editing via Tangent-Space Guidance and Optimization

TanGO:通过切线空间引导和优化实现免训练3D编辑

Siwoo Lim, Sunjae Yoon, Gwanhyeong Koo, Hyeonseo Yun, Chang D. Yoo

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院) Chung-Ang University(中央大学)

AI总结 研究针对3D生成模型免训练编辑存在语义伪影的问题,提出TanGO框架,通过在切线空间自适应逐令牌引导,制定最优控制规则并确定控制信号强度,减少结构伪影,性能优于现有基线。

Comments ECCV 2026

详情
AI中文摘要

近期基于流匹配的3D生成模型(如VecSet)采用结构化表示,但令牌共享全局上下文,导致传统免训练编辑存在语义伪影。为此提出TanGO,一个免训练框架,能在生成动力学的切线空间中实现自适应逐令牌引导。通过制定单步最优控制规则,并利用从源和目标速度场导出的冯·米塞斯 - 费舍尔启发的方向差异确定每个令牌控制信号的强度。实验表明TanGO显著减少结构伪影并实现了优于现有3D编辑基线的性能。

英文摘要

While recent flow-matching 3D generative models (e.g., VecSet) adopt structured representations, their tokens share global context, causing conventional training-free editing to suffer from semantic artifacts such as collapsed preserved regions or incomplete transformations. To address this, we propose TanGO, a training-free framework that enables adaptive per-token steering in the tangent space of generative dynamics. To realize this selective control, we formulate a one-step optimal control rule and determine the strength of each token's control signal using a von Mises-Fisher inspired directional discrepancy derived from the source and target velocity fields. Experiments show that TanGO substantially reduces structural artifacts and achieves state-of-the-art performance, outperforming existing 3D editing baselines. The code is publicly available at https://github.com/siw00-lim/TanGO.

URL PDF HTML 收藏
2607.03454 2026-07-17 cs.RO cs.LG 版本更新

ADP: Adversarial Dynamics Priors for Physically Grounded Humanoid Locomotion

ADP:用于物理基础类人机器人运动的对抗动力学先验

Seokju Lee, Jeongtae Lee, Jeonghyeok Lim, Jeonguk Kang, Byungwook Lee, Seungho Han, Keun Ha Choi, Dongil Park, Kyung-Soo Kim

机构 * Mechatronics, Systems and Control Lab (MSC Lab), Department of Mechanical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院机械工程系机电一体化、系统与控制实验室(MSC实验室)) Samsung Electronics, Future Robotics AI Group(三星电子未来机器人人工智能集团) School of Electrical Engineering, Hanyang University(汉阳大学电气工程学院) Advanced Robotics Research Center, Korea Institute of Machinery & Materials (KIMM)(韩国机械与材料研究所先进机器人研究中心)

AI总结 提出对抗动力学先验(ADP)用于类人机器人抗干扰运动控制。以动力学特征取代运动学特征作对抗目标,用轨迹优化建参考数据集训练鉴别器,提升机器人抗干扰及运动跟踪能力。

Comments 8 pages, 6 figures

详情
AI中文摘要

本文中,我们提出了用于抗干扰类人机器人运动控制的对抗动力学先验(ADP)。现有的基于运动先验的方法通过模仿运动学运动特征来诱导自然运动风格,但它们没有直接对动力学特征进行正则化,如质心运动、质心动量、接触力和接触状态。为了解决这一限制,我们用从运动轨迹中提取的选定动力学特征取代运动学运动风格特征作为对抗目标。为此,我们使用轨迹优化来构建一个参考数据集,并训练一个鉴别器来评估策略诱导的时间窗口是否与所得参考一致。通过显式运动跟踪,ADP鼓励策略展开即使在受到干扰后也保持接近参考支持。实验结果表明,与我们评估中最强的基线AMP相比,ADP将80%成功脉冲阈值($J_{80}$)提高了16.7%,同时将方向平均恢复时间和速度跟踪误差分别降低了47.9%和35.4%。

英文摘要

In this paper, we propose Adversarial Dynamics Priors (ADP) for perturbation-resilient humanoid locomotion control. Existing motion prior-based methods induce natural motion styles by imitating kinematic motion features, but they do not directly regularize dynamics features, such as CoM motion, centroidal momentum, contact forces, and contact states. To address this limitation, we replace kinematic motion-style feature with selected dynamics features extracted from locomotion trajectories as the target of adversarial regularization. To this end, we use trajectory optimization to construct a reference dataset and train a discriminator to evaluate whether policy-induced temporal windows are consistent with the resulting reference distribution. Without explicit motion tracking, ADP encourages policy rollouts to remain close to the reference support, even after perturbations. Experimental results show that, compared with AMP, the strongest baseline in our evaluation, ADP improves the $80\%$-success impulse threshold ($J_{80}$) by $16.7\%$, while reducing direction-averaged recovery time and velocity tracking error by $47.9\%$ and $35.4\%$, respectively.

URL PDF HTML 收藏
2503.19877 2026-07-17 cs.CL 版本更新

Scaling Evaluation-time Compute with Reasoning Models as Evaluators

通过推理模型作为评估器来提升评估时的计算能力

Seungone Kim, Ian Wu, Jinu Lee, Xiang Yue, Seongyun Lee, Mingyeong Moon, Carolin Lawrence, Kiril Gashteovski, Julia Hockenmaier, Graham Neubig, Sean Welleck

机构 * CMU(卡内基梅隆大学) UIUC(伊利诺伊大学) KAIST AI(韩国科学技术院人工智能研究所) NEC Laboratories Europe(日本 NEC 欧洲实验室) Ss.Cyril and Methodius University of Skopje(斯科普里塞尔吉尔和梅蒂乌斯大学)

AI总结 本文探讨了通过增加评估时的计算量来提升语言模型的评估能力,利用推理模型作为评估器,分别评估响应整体和每个步骤,从而提高评估效果。

Comments ACL 2026 Findings

详情
AI中文摘要

随着语言模型(LM)的输出越来越自然,评估其质量变得越来越困难。同时,通过增加测试时的计算量来提升LM的'思考'时间,已被证明是解决数学和代码等领域挑战性问题的有效技术。这引发了一个自然的问题:是否可以通过增加测试时的计算量来提升LM的评估能力?为回答这个问题,我们研究了利用推理模型——即能够原生生成长链推理的LM——作为评估器。具体而言,我们考察了通过(1)使用推理模型,以及(2)提示这些模型不仅评估响应整体(即结果评估),还评估响应中的每个步骤(即过程评估)来利用更多测试时计算量的方法。在实验中,我们观察到评估器的性能随着生成更多推理标记而单调提升,类似于LM生成中的趋势。此外,我们使用这些更准确的评估器对多个生成进行重新排序,并证明在评估时花费更多计算量可以像在生成时花费更多计算量一样有效,从而提升LM的问题解决能力。

英文摘要

As language model (LM) outputs get more and more natural, it is becoming more difficult than ever to evaluate their quality. Simultaneously, increasing LMs' "thinking" time through scaling test-time compute has proven an effective technique to solve challenging problems in domains such as math and code. This raises a natural question: can an LM's evaluation capability also be improved by spending more test-time compute? To answer this, we investigate employing reasoning models-LMs that natively generate long chain-of-thought reasoning-as evaluators. Specifically, we examine methods to leverage more test-time compute by (1) using reasoning models, and (2) prompting these models to evaluate not only the response as a whole (i.e., outcome evaluation) but also assess each step in the response separately (i.e., process evaluation). In experiments, we observe that the evaluator's performance improves monotonically when generating more reasoning tokens, similar to the trends observed in LM-based generation. Furthermore, we use these more accurate evaluators to rerank multiple generations, and demonstrate that spending more compute at evaluation time can be as effective as using more compute at generation time in improving an LM's problem-solving capability.

URL PDF HTML 收藏
2607.13579 2026-07-16 cs.RO cs.AI cs.LG 新提交

Agile perceptive multi-skill locomotion for quadrupedal robots in the wild

野外四足机器人的敏捷感知多技能运动

Jun-Gill Kang, Jaehyun Park, Tae-Gyu Song, Joon-Ha Kim, Seungwoo Hong, Hae-Won Park

机构 * Agency for Defense Development(国防发展局) Korea Advanced Institute of Science and Technology(韩国科学技术院) DIDEN Robotics(迪登机器人公司) Korea University(韩国大学)

AI总结 研究使四足机器人在复杂地形实现多技能运动的问题,提出 APT-RL 框架,利用机载感知和计算自主转换技能,通过轨迹优化生成数据集训练技能,经实验验证该框架能让机器人在复杂环境敏捷机动,稳健穿越多样障碍物。

Comments Project page: https://skillquadsr.github.io/ ,This is the author's version of the work. It is posted here by permission of the AAAS for personal use, not for redistribution. The definitive version was published in Science Robotics on 7.15.2026; doi: 10.1126/scirobotics.adz7397. Jun-Gill Kang and Jaehyun Park are co-first authors. Seungwoo Hong and Hae-Won Park are co-corresponding authors

详情
AI中文摘要

使四足机器人穿越复杂地形(从崎岖的户外环境到城市景观),需要多种运动技能的无缝集成、步态间的平滑过渡以及仅使用机载传感器的高速感知运动。我们提出了 APT-RL(基于动作预训练变压器的强化学习),这是一个统一框架,通过仅利用机载感知和计算的自主技能转换,实现多技能运动以在复杂环境中高速穿越。我们的方法通过简化动力学的轨迹优化生成大规模、特征丰富的 2D 运动数据集。这些数据集能训练多样、可复用的运动技能,有效转移到在复杂不平地形上运行的真实四足机器人。高质量技能为高效学习复杂下游任务提供强先验,并自然扩展到 3D 环境,实现部署策略中的平滑、高速多技能运动。实际实验证明了该框架的能力:机器人能通过复杂室内障碍物和户外野外环境进行敏捷机动,包括达到每秒 6 米瞬时峰值速度的动态下拉机动。单一机载策略能稳健穿越各种障碍物,展示了我们方法的通用性和有效性。

英文摘要

Enabling quadrupedal robots to traverse complex terrains-from rugged outdoor environments to urban landscapes-requires seamless integration of multiple motor skills, smooth transitions between gaits, and high-speed perceptive locomotion using only onboard sensors. We present APT-RL (Action Pretrained Transformer-based Reinforcement Learning), a unified framework that enables multi-skill locomotion to achieve high-speed traversal in complex environments through autonomous skill transitions utilizing only onboard perception and computation. Our approach generates large-scale, feature-rich 2D motion datasets through trajectory optimization with simplified dynamics. These datasets enable training of diverse, reusable locomotion skills that transfer effectively to a real quadruped robot operating on complex uneven terrains. The resulting high-quality skills serve as strong priors for efficient learning of complex downstream tasks and extend naturally to 3D environments, enabling smooth, high-speed multi-skill locomotion in deployed policy. Real-world experiments demonstrate the framework's capabilities: the robot performs agile maneuvers through complex indoor obstacles and outdoor wild environments, including dynamic drop-down maneuvers that reach instantaneous peak speeds of up to 6 meters per second. A single onboard policy enables robust traversal of diverse obstacles, including stairs, hurdles, stepping stones, gaps, and fallen branches, demonstrating the versatility and effectiveness of our approach.

URL PDF HTML 收藏
2607.08803 2026-07-16 q-bio.QM cs.AI cs.LG 版本更新

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

TheBioCollection:用于生物学的统一预训练规模语言模型语料库

Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung

机构 * Trillion Labs KAIST(韩国科学技术院) SK Biopharmaceuticals Co., Ltd.(SK生物制药公司) Lunit Inc.(Lunit公司) AIGEN Sciences Inc.(AIGEN科学公司)

AI总结 为满足生物学大语言模型训练需求,提出TheBioCollection语料库,整合异构生物资源并丰富记录、引入新任务,配对TheBioCollection-Eval评估,固定架构训练后模型在各领域表现提升,语言能力基本不变。

详情
AI中文摘要

向生物学大语言模型(BioLM)的发展产生了对训练语料库的需求,以便赋予模型对生物学的真正理解。然而,现有的生物资源分散在异构格式中,未被组织成用于语言模型训练的连贯语料库。我们提出了TheBioCollection,一个526亿token的预训练规模语料库,将这些不同资源转换为统一的、可用于训练的形式。它不仅整合现有数据,还用工具计算的生物学特性丰富每条记录,并引入新的指令任务。我们将该语料库与TheBioCollection-Eval配对进行评估。在固定基础Gravity-16B-A3B架构的情况下,在TheBioCollection上训练使模型在TheBioCollection-Eval上的总分提高了一倍多,且各领域均有提升,同时基本保持一般语言能力不变。

英文摘要

The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with TheBioCollection-Eval, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on TheBioCollection more than doubles its overall score on TheBioCollection-Eval with gains in every domain, while leaving general linguistic ability nearly intact.

URL PDF HTML 收藏
2607.01060 2026-07-16 cs.RO 版本更新

RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation

RoboWorld: 用于通用机器人策略评估的快速可靠神经模拟器

Byeongguk Jeon, Seonghyeon Ye, JaeHyeok Doo, Sungdong Kim, Minjoon Seo, Hyungmok Son, Kimin Lee

机构 * KAIST(韩国科学技术院) Config

AI总结 提出RoboWorld自动化评估流程,结合快速自回归视频世界模型和任务进度感知视觉语言模型评分,通过Step Forcing减少训练-测试不匹配,实现与真实世界评估高度一致。

Comments Project page: https://byeongguks.github.io/RoboWorld/

详情
AI中文摘要

视频世界模型正成为评估通用机器人策略的可扩展替代方案,绕过了真实世界部署的物理限制和工程负担。然而,使用视频世界模型评估策略仍然具有挑战性,因为世界模型误差可能使生成的轨迹不可靠,且推理速度慢限制了大规模吞吐量。我们引入了RoboWorld,一种自动化评估流程,将快速自回归视频世界模型与任务进度感知的视觉语言模型评分相结合。为了实现可靠的长程自回归世界模型轨迹生成,我们提出了Step Forcing,它结合了锚定和单步自前向上下文,以减少训练-测试不匹配,同时保留动作-观察动态。这些组件共同使RoboWorld能够在不同任务和环境中与真实世界机器人评估高度一致,达到Pearson's r = 0.989和Spearman's ρ = 0.970。

英文摘要

Video world models are emerging as a scalable alternative for evaluating generalist robot policies, bypassing the physical constraints and engineering burdens of real-world deployment. However, evaluating policies with video world models remains challenging, as world-model errors can make generated rollouts unreliable and slow inference limits large-scale throughput. We introduce RoboWorld, an automated evaluation pipeline that pairs a fast autoregressive video world model with a task-progress-aware vision-language model scoring. To enable reliable long-horizon autoregressive world-model rollouts, we propose Step Forcing, which combines anchored and one-step self-forwarded contexts to reduce train-test mismatch while preserving action-observation dynamics. Together, these components enable RoboWorld to align strongly with real-world robot evaluation across tasks and environments, achieving Pearson's r = 0.989 and Spearman's $ρ$ = 0.970.

URL PDF HTML 收藏
2607.12829 2026-07-15 cs.LG cs.AI cs.CL 新提交

Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques

加速掩码扩散大语言模型:高效推理技术综述

Daehoon Gwak, Minhyung Lee, Junwoo Park, Jaegul Choo

机构 * KAIST AI(韩国科学技术院人工智能研究所) Yonsei University(延世大学)

AI总结 综述介绍用于扩散大语言模型的统一延迟分解框架,以理清算法、架构和系统因素对推理速度的影响,将加速技术分类,提供可重复基准测试指导方针并强调实现并行生成潜力的挑战。

Comments Accepted at IJCAI-ECAI 2026 (Survey Track)

详情
AI中文摘要

扩散大语言模型(dLLMs)在并行生成方面相对于标准自回归模型具有理论优势。然而,仅并行生成并不能保证实际加速。实现这种效率需要专门的推理机制,如扩散感知缓存和重用。随着推理效率成为实际部署的先决条件,近期研究积极探索跨算法、架构和系统的加速技术。但由于现有基准测试中算法、架构和系统级因素之间复杂的权衡导致端到端延迟难以进行严格比较。本综述引入统一延迟分解框架来理清这些因素并分析其对实际部署中推理速度的影响。在此框架指导下,将加速技术沿算法创新、架构与系统优化以及推理时间缩放三个轴进行分类。最后提供可重复基准测试的指导方针并强调实现并行生成全部潜力的开放挑战。

英文摘要

Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency requires specialized inference mechanisms, such as diffusion-aware caching and reuse. Consequently, as inference efficiency becomes a prerequisite for practical deployment, recent research has actively explored acceleration techniques across algorithms, architectures, and systems. However, rigorous comparisons remain difficult, as end-to-end latency stems from intricate trade-offs between algorithmic, architectural, and system-level factors that are often conflated in existing benchmarks. In this survey, we introduce a unified latency decomposition framework for dLLMs to disentangle these factors and analyze their impact on inference speed in real deployments. Guided by this framework, we categorize acceleration techniques along three axes covering algorithmic innovations, architectural and system optimizations, and inference-time scaling. Finally, we provide guidelines for reproducible benchmarking and highlight open challenges for realizing the full potential of parallel generation.

URL PDF HTML 收藏
2605.22432 2026-07-15 cs.LG 版本更新

AMUSE: Anytime Muon with Stable Gradient Evaluation

AMUSE: 任何时刻的Muon with Stable Gradient Evaluation

Jueun Kim, Baekrok Shin, Jihun Yun, Beomhan Baek, Minhak Song, Chulhee Yun

机构 * KAIST(韩国科学技术院) KRAFTON(KRAFTON公司) Seoul National University(首尔国立大学)

AI总结 本文研究了Muon算法的机制,提出了一种名为AMUSE的算法,通过结合Muon的快速批量进步和Schedule-Free平均的稳定效果,实现了无需学习率调度的任何时刻训练,并在视觉任务和大语言模型预训练中提升了性能-迭代帕累托前沿。

Comments 44 pages, 27 figures

详情
AI中文摘要

现代深度学习通常依赖于AdamW和预设的学习率调度,但最近的研究挑战了这两个组件:Schedule-Free优化通过迭代平均去除显式调度,而Muon通过正交化动量来改进矩阵参数的更新几何。尽管Muon在经验上表现强劲,但其底层机制仍部分不明确。我们通过河谷损失景观研究Muon,其中有用的训练进展发生在平坦、低曲率的 bulk 子空间(河流)中,而高曲率主导方向形成陡峭的河谷墙壁,导致振荡。我们实证显示,Muon的正交化通过增加bulk成分加速河流进展,但也放大了主导方向的噪声,导致振荡轨迹。基于此,我们提出Anytime MUon with Stable gradient Evaluation (AMUSE),它结合Muon的快速bulk进展与Schedule-Free平均的稳定效果。AMUSE使用一个随时间变化的插值系数,最初评估接近快速Muon序列的梯度以实现快速适应,然后逐渐转向稳定的平均序列以抑制河谷墙壁的振荡。结果,AMUSE不需要学习率调度并支持任何时刻训练。在视觉任务和大语言模型预训练中,AMUSE在性能-迭代帕累托前沿上一致优于(Schedule-Free) AdamW和Muon。

英文摘要

Modern deep learning commonly relies on AdamW with prescribed learning rate schedules, but recent works challenge both components: Schedule-Free optimization removes explicit schedules via iterate averaging, and Muon improves the update geometry by orthogonalizing momentum for matrix parameters. Despite Muon's strong empirical performance, its underlying mechanism remains partially understood. We study Muon through the river-valley loss landscape, where useful training progress occurs along a flat, low-curvature bulk subspace (the river), while high-curvature dominant directions form steep valley walls that induce oscillations. We empirically show that while Muon's orthogonalization accelerates river progress by increasing the bulk component, it also amplifies dominant-direction noise, causing oscillatory trajectories. Building on this, we propose Anytime MUon with Stable gradient Evaluation (AMUSE), which integrates Muon's rapid bulk progress with the stabilizing effect of Schedule-Free averaging. AMUSE uses a time-varying interpolation coefficient that initially evaluates gradients near the fast Muon sequence for rapid adaptation, then gradually shifts toward the stable averaged sequence to suppress valley-wall oscillations. As a result, AMUSE requires no learning rate schedules and supports anytime training. Across vision tasks and large language model pretraining, AMUSE consistently improves the performance-iteration Pareto frontier over (Schedule-Free) AdamW and Muon.

URL PDF HTML 收藏
2510.00492 2026-07-15 cs.AI 版本更新

Rethinking Reward Models for Multi-Domain Test-Time Scaling

重新思考多领域测试时扩展的奖励模型

Dong Bok Lee, Seanie Lee, Sangwoo Park, Minki Kang, Jinheon Baek, Dongki Kim, Dominik Wagner, Jiongdao Jin, Heejun Lee, Tobias Bocklet, Jinyu Wang, Jingjing Fu, Sung Ju Hwang, Jiang Bian, Lei Song

机构 * KAIST(韩国科学技术院) Microsoft Research Asia(微软亚洲研究院) TH Nürnberg

AI总结 研究多领域测试时扩展的奖励模型,对四种奖励模型变体进行统一评估,发现dORM与dPRM相当,gPRM无竞争力,gORM最稳健,挑战细粒度监督更好的假设,支持生成式结果验证并公开代码。

详情
AI中文摘要

大语言模型在测试时扩展期间的可靠性通常用区分正确推理和有缺陷逻辑的外部验证器或奖励模型来评估。先前工作研究了仅评估最终答案的结果奖励模型(ORM)和对中间推理步骤评分的过程奖励模型(PRM)。尽管PRM常因更细粒度监督被视为有利,但支持证据多来自数学相关设置,其在更广泛领域的相对优势仍不明。本文对四种奖励模型变体在14个不同领域进行首次统一评估,发现dORM与dPRM表现相当,gPRM无竞争力,gORM最稳健。gPRM性能差归因于逐步评分过程继承基于大语言模型自动标注的标签噪声。研究结果挑战了细粒度监督总是更好的假设,支持多领域部署的生成式结果验证,代码公开。

英文摘要

The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward models} that distinguish correct reasoning from flawed logic. Prior work has studied both outcome reward models (ORMs), which assess only the final answer, and process reward models (PRMs), which score intermediate reasoning steps. Although PRMs are often viewed as advantageous due to their finer-grained supervision, much of the supporting evidence comes from math-adjacent settings, and their relative benefits across broader domains remain unclear. We present the first unified evaluation of four reward model variants, discriminative ORM and PRM (dORM, dPRM) and generative ORM and PRM (gORM, gPRM), across 14 diverse domains. Contrary to conventional wisdom, we find that (i) dORM performs on par with dPRM, (ii) gPRM is not competitive, and (iii) overall, gORM is the most robust, yielding significant and consistent gains across every tested domain. We attribute the worse performance of gPRM to the stepwise scoring process, which inherits label noise from LLM-based automatic labeling, leading to difficulties in evaluating long reasoning trajectories, including those involving self-correcting reasoning. Both our theoretical analysis and empirical observations indicate that stepwise aggregation compounds errors as reasoning length increases. These findings challenge the common assumption that fine-grained supervision is always better and support generative outcome verification for multi-domain deployment. Our \href{https://github.com/db-Lee/Multi-RM}{\underline{code}} is publicly available to facilitate future research in multi-domain settings.

URL PDF HTML 收藏
2607.11498 2026-07-14 cs.RO cs.AI 新提交

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

像机器人一样看:用于视觉-语言-动作模型的以机器人为中心的点图

Byungkun Lee, Dongyoon Hwang, Dongjin Kim, Hojoon Lee, Minho Park, Jaegul Choo

机构 * KAIST AI(韩国科学技术院人工智能研究所) Holiday Robotics(假日机器人公司)

AI总结 研究视觉-语言-动作模型中因观察与动作定义的帧不匹配问题,提出以机器人为中心的点图方法,该方法能保留2D VLA所需网格,以最小架构变化集成到现有模型,实验证明其在模拟和实际机器人实验中表现优于基线。

Comments Project page: https://davian-robotics.github.io/pointmap/

详情
AI中文摘要

视觉-语言-动作(VLA)模型根据视觉观察和语言指令预测机器人动作。动作在机器人自身的3D坐标系中定义,但大多数VLA在相机帧中观察场景,导致观察场景的位置与定义动作的位置之间存在帧不匹配。在固定视点下这种不匹配影响较小,而随着大规模数据集聚合不同相机设置下的演示,且策略必须跨视点进行泛化时,问题变得更严重。我们以机器人为中心的点图来解决这种不匹配,其像素存储机器人帧中场景点的3D坐标。点图提供机器人帧3D几何结构,同时保留预训练2D VLA所需的密集H x W网格,能以最小架构变化集成到现有VLA中。在RoboCasa上,点图改进了pi0.5和SmolVLA,优于代表性的相机视点和3D感知基线。在实际机器人实验中,当相机移到训练中未见过的位置时,相对于仅使用RGB的策略,点图优势更明显。

英文摘要

Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene is observed and where actions are defined. The mismatch is benign under a fixed viewpoint, where the policy can memorize a single observation-to-action mapping, but grows harder as large-scale datasets aggregate demonstrations across diverse camera setups and the policy must generalize this mapping across viewpoints. We address this mismatch with robot-centric pointmaps, images whose pixels store the 3D coordinates of scene points in the robot frame. Pointmaps provide robot-frame 3D geometry while preserving the dense H x W grid expected by pretrained 2D VLAs, so they integrate into existing VLAs with minimal architectural change. On RoboCasa, pointmaps improve both pi0.5 and SmolVLA and outperform representative camera-viewpoint and 3D-aware baselines. In real-robot experiments, their advantage over an RGB-only policy widens when the camera is moved to a placement unseen during training.

URL PDF HTML 收藏
2607.11041 2026-07-14 cs.RO 新提交

PAKE: Learning Whole-Body Loco-Manipulation with Partial Kinematic Embeddings

PAKE:利用部分运动学嵌入学习全身移动操作

Zhengmao He, Moonkyu Jung, Hyeongjun Kim, Jiseong Lee, Hui Zhang, Jemin Hwangbo, Jie Song

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Korea Advanced Institute of Science and Technology(韩国科学技术院) ETH Zurich(苏黎世联邦理工学院)

AI总结 研究针对高自由度机器人全身移动操作面临的挑战,提出将其分解为部分参考运动生成和低级模仿控制的框架,用KNF模型生成参考运动,经高低级控制器实现精确控制,在模拟和硬件实验中均表现优异,为相关操作提供实用方案。

详情
AI中文摘要

移动操作展现出了很有前景的能力。然而,实现高精度控制、管理由多自由度引发的高维动作空间以及充分利用全身系统的固有冗余仍具有挑战性。本文提出了一种新颖的全身控制框架,通过将复杂的移动操作问题分解为部分参考运动生成和低级模仿控制来有效应对这些挑战。引入了一种在大规模运动学数据集上训练的新运动学归一化流(KNF)模型来生成多样且可行的部分参考运动。训练了高级控制器在KNF的潜在空间中导航以利用冗余解,低级控制器确保物理上可行且精确的运动执行。在配备六自由度机械臂的四足机器人上验证了该方法。模拟实验结果表明该方法在跟踪精度和可行工作空间覆盖方面显著优于现有方法。硬件部署评估中,系统在8种不同的移动操作任务的24个情节上实现了末端执行器姿态跟踪误差为4.5厘米和0.14弧度,同时分别以0.1米/秒和0.01弧度/秒的线性和角速度误差保持精确的运动跟踪,优于竞争基线。我们的方法为高自由度机器人系统中的精确和通用全身移动操作提供了实用且强大的解决方案,对各种下游机器人任务具有潜在的应用前景。

英文摘要

Loco-manipulation has recently shown promising capabilities; however, achieving high-precision control, managing the high-dimensional action space induced by many degrees of freedom (DoFs), and fully exploiting the inherent redundancy of whole-body systems remain challenging. In this paper, we propose a novel whole-body control framework that effectively addresses these challenges by decomposing the complex loco-manipulation problem into partial reference motion generation and low-level imitation control. We introduce a new Kinematic Normalizing Flow (KNF) model, trained on a large-scale kinematic dataset, that generates diverse yet feasible partial reference motions. A high-level controller is then trained to navigate the KNF's latent space to exploit redundant solutions, while a low-level controller ensures physically feasible and accurate motion execution. We validate our approach on the quadrupedal robot equipped with a six-DoF robotic arm. In simulation, experimental results show that our approach significantly outperforms state-of-the-art methods in terms of tracking accuracy and feasible workspace coverage. For hardware deployment, we evaluate the system over 24 episodes across 8 different mobile loco-manipulation tasks. The system achieves end-effector pose-tracking errors of 4.5 cm and 0.14 rad, while maintaining accurate locomotion tracking with linear and angular velocity errors of 0.1 m/s and 0.01 rad/s, respectively, outperforming competitive baselines. Our method represents a practical and powerful solution for accurate and generalized whole-body loco-manipulation in high-DoF robotic systems, with promising potential for diverse downstream robotic tasks.

URL PDF HTML 收藏
2607.11031 2026-07-14 cs.RO 新提交

GraspGraphNet: Graph-Structured Multi-Embodiment Dexterous Grasp Generation

GraspGraphNet:基于图结构的多机器人灵巧抓取生成

Yeonseo Lee, Taeyeop Lee, Hyosup Shin, Guebin Hwang, Sungho Jo

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

AI总结 研究跨机器人手的灵巧抓取生成难题,提出GraspGraphNet框架,将手表示为运动学图,结合多种技术建模交互,直接在相关空间应用条件流匹配,无需后处理等,共享模型在多场景取得高成功率,证明图结构手部表示的有效性。

Comments Project: https://lysees.github.io/graspgraphnet-page

详情
AI中文摘要

跨机器人手的灵巧抓取生成具有挑战性,因为手在运动拓扑、驱动维度和原生命令空间方面存在差异。我们引入了GraspGraphNet,这是一个拓扑感知的抓取生成框架,它将每只手表示为从URDF派生的运动学图,并直接生成可执行的手掌姿势和关节配置。GraspGraphNet结合了分层物体表面编码、可微正向运动学和动态世界边缘消息传递,以对不断演变的机器人-物体交互进行建模。它直接在可执行的手掌姿势和关节状态空间中应用条件流匹配,避免了后处理优化、逆运动学和重新定位。使用在巴雷特手、阿莱格罗手和影子手上训练的共享模型,GraspGraphNet在40个物体的基准测试中,每次抓取的推理时间为40毫秒,平均成功率达到83.48%。在不重新训练的情况下,同一模型在受控手指移除变体上的成功率为72.70%,证明了对手部拓扑变化的鲁棒性。这些结果表明,图结构的手部表示可以有效地支持具有不同运动结构的机器人手的灵巧抓取生成。

英文摘要

Dexterous grasp generation across robot hands is challenging because hands differ in kinematic topology, actuation dimensions, and native command spaces. We introduce GraspGraphNet, a topology-aware grasp generation framework that represents each hand as a URDF-derived kinematic graph and directly generates executable palm poses and joint configurations. GraspGraphNet combines hierarchical object surface encoding, differentiable forward kinematics, and dynamic world-edge message passing to model evolving robot-object interactions. It applies conditional flow matching directly in executable palm-pose and joint-state space, avoiding post-processing optimization, inverse kinematics, and retargeting. Using a shared model trained on Barrett Hand, Allegro Hand, and Shadow Hand, GraspGraphNet achieves an average success rate of 83.48% with 40ms inference time per grasp on a 40-object benchmark. Without retraining, the same model achieves 72.70% success on controlled finger-removal variants, demonstrating robustness to hand-topology variations. These results suggest that graph-structured hand representations can effectively support dexterous grasp generation across robot hands with different kinematic structures. Project: https://lysees.github.io/graspgraphnet-page

URL PDF HTML 收藏
2607.09753 2026-07-14 cs.CV cs.AI cs.LG 新提交

Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis

通过内部潜变量分析对扩散模型进行统一骨干细化

Haksoo Lim, Myeongjin Lee, Wonjoon Chang, Jaesik Choi

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) INEEJI

AI总结 研究扩散模型骨干细化问题,提出DUNE框架,通过分析内部潜变量检测突变偏差,对选定条目进行骨干特定抑制,可自然扩展到基于Transformer的模型,经实验验证该方法能提高保真度并减少幻觉。

Comments 45 pages, 23 figures. Accepted at the European Conference on Computer Vision (ECCV) 2026

详情
AI中文摘要

扩散模型在各个领域取得了显著成功,其性能与参数化得分函数的去噪骨干密切相关。本文对扩散组件进行了系统的、阶段感知分析,发现深度潜变量中的早期突变与伪像密切相关。基于此,引入了DUNE(扩散统一网络细化器),这是一个无需训练的细化框架,利用基于共享EMA的准则检测深度低噪声内部潜变量中的突变偏差,并对检测器选择的条目应用特定骨干的抑制。该原理可自然扩展到基于Transformer的扩散模型。大量实验表明,DUNE提高了保真度并减少了幻觉。

英文摘要

Diffusion models have achieved remarkable success across diverse domains, with performance closely related to the denoising backbones that parameterize the score function. In this paper, we present a systematic, phase-aware analysis of diffusion components and show that abrupt, early-stage fluctuations in deep latents are strongly associated with artifacts. Guided by these findings, we introduce DUNE (Diffusion Unified Network refiNEr), a training-free refinement framework that detects abrupt deviations in deep low-noise internal latents using a shared EMA-based criterion, and applies backbone-specific suppression to the detector-selected entries. Although derived from U-Net, the same detect-suppress principle extends naturally to Transformer-based diffusion models by acting on the latents of deep self-attention blocks. Extensive experiments across multiple backbones indicate that DUNE improves fidelity while reducing hallucinations, offering new insight into where and when diffusion backbones should be controlled.

URL PDF HTML 收藏
2606.06065 2026-07-14 cs.CL cs.SD eess.AS 版本更新

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition

多任务学习还不够:双输出第二语言语音识别中的表示纠缠

Seung Hwan Cho, Young-Min Kim

机构 * KAIST(韩国科学技术院)

AI总结 针对双输出第二语言语音识别,研究发现多任务学习导致表面转录性能下降,归因于编码器级别的表示纠缠,尤其在英语中随表面-意义差异增大而加剧。

Comments 5 pages, 2 figures, Accepted to the 43rd International Conference on Machine Learning Workshop on Machine Learning for Audio

详情
AI中文摘要

第二语言(L2)语音识别通常需要发音转录和预期意义的转录。多任务学习(MTL)是一种自然的方法,因为它假设共享表示对两个输出都有益。然而,本文表明这一假设在韩语和英语中并不成立。MTL提高了意义转录但降低了表面转录,尤其是在英语中,性能下降与通过Levenshtein编辑距离测量的表面-意义差异成正比。编码器分析将这些模式与编码器级别的纠缠联系起来,韩语保留了不同的任务表示,而英语产生了几乎相同的表示。跨任务解码器分析表明,意义双输出解码器适应了独特的表示,而表面双输出解码器仍受编码器约束。这些发现促使设计能够减轻编码器级别纠缠的MTL框架,以减少双输出L2自动语音识别中的表面性能下降。

英文摘要

Second-language (L2) speech recognition often requires transcriptions of pronunciations and intended meanings. Multi-task learning (MTL) is a natural approach because it assumes that shared representations benefit both outputs. However, this paper shows that this assumption does not hold across Korean and English. MTL improves meaning but degrades surface transcription, especially in English, where the degradation scales with surface-meaning divergence measured by Levenshtein edit distance. Encoder analysis links these patterns to encoder-level entanglement, with Korean preserving disentangled representations while English produces nearly identical ones. Cross-output decoder analysis shows that the meaning dual-output decoder adapts with a unique representation, while the surface dual-output decoder remains constrained by the encoder. These findings motivate the design of MTL frameworks that mitigate encoder-level entanglement to reduce surface degradation in dual-output L2 automatic speech recognition.

URL PDF HTML 收藏
2605.11504 2026-07-14 cs.LG cs.CR 版本更新

CTFusion: A CTF-based Benchmark for LLM Agent Evaluation

CTFusion: 一种基于CTF的LLM代理评估基准

Dongjun Lee, Ga-eun Bae, Insu Yun

机构 * School of Electrical Engineering, KAIST(韩国科学技术院电子工程学院)

AI总结 本文提出CTFusion,一种基于CTF的评估框架,通过减少竞争影响和数据污染,提升对LLM代理的评估可靠性。

Comments Accepted at the Second Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD), ICML 2026. Revised to match the camera-ready version. OpenReview: https://openreview.net/forum?id=aUSUvS4dPL

详情
AI中文摘要

近年来,大语言模型(LLM)的进步使复杂多步骤任务的代理系统成为可能,网络安全成为重要应用。为评估此类代理,研究者广泛采用捕获旗帜(CTF)基准。然而,现有CTF基准复用现有挑战,导致数据污染和潜在作弊。我们通过集成网络搜索工具验证了这些问题。为解决这些限制,我们提出CTFusion,一种基于实时CTF的流评估框架。CTFusion在单个团队账户下保持代理独立性,并通过仅转发每个挑战的第一个正确旗帜来减少竞争影响。此外,我们实现了CTFusion作为模型上下文协议(MCP)服务器,适用于广泛使用的CTFd平台,适用于多样化的CTF事件和代理类型。通过实验验证,现有CTF基准在评估LLM代理时不可靠,而CTFusion能作为评估网络安全代理的稳健解决方案。我们开源CTFusion以促进该领域的未来研究。

英文摘要

Recent advances in Large Language Models (LLMs) have enabled agentic systems for complex, multi-step tasks; cybersecurity is emerging as a prominent application. To evaluate such agents, researchers widely adopt Capture The Flag (CTF) benchmarks. However, current CTF benchmarks reuse existing challenges, which exposes them to data contamination and potential cheating. Notably, we confirmed these issues in practice by integrating web search tools into an existing agent. To address these limitations, we present CTFusion, a streaming evaluation framework built on Live CTFs. To achieve this, CTFusion preserves per-agent independence under a single team account and reduces competition impact by forwarding only the first correct flag per challenge. Moreover, we implement CTFusion as a Model Context Protocol (MCP) server on the widely used CTFd platform, which offers broad applicability to diverse CTF events and agent types. Through experiments with three LLMs, two agents, and five Live CTFs, we demonstrate that existing CTF benchmarks can be unreliable in assessing LLM-based agents, while CTFusion can serve as a robust solution for evaluating cybersecurity agents. We release CTFusion as open source to foster future research in this area.

URL PDF HTML 收藏
2604.21334 2026-07-14 cs.AI cs.CE cs.CL cs.LG econ.GN q-fin.EC 版本更新

Ideological Bias in LLMs' Economic Causal Reasoning

大语言模型在经济因果推理中的意识形态偏见

Donggyu Lee, Hyeok Yun, Jungwon Kim, Junsik Min, Sungwon Park, Sangyoon Park, Jihee Kim

机构 * Graduate School of Data Science, KAIST(韩国科学技术院数据科学研究生院) College of Business, KAIST(韩国科学技术院商学院) School of Computing, KAIST(韩国科学技术院计算机学院) Division of Social Science, HKUST(香港科技大学社会科学系)

AI总结 研究探讨LLM在经济因果推理中的意识形态偏见,通过扩展EconCausal基准测试,评估20种先进LLM在处理意识形态冲突案例时的准确性,发现LLM在意识形态冲突案例上表现更差,且存在系统性偏倚。

Comments Accepted at COLM 2026

详情
AI中文摘要

大语言模型(LLM)在推理经济因果效应时是否存在系统性的意识形态偏见?随着LLM在政策分析和经济报告中的广泛应用,正确判断因果方向至关重要。本文通过扩展EconCausal基准测试,引入意识形态冲突案例——干预导向(支持政府)和市场导向(支持市场)视角预测因果方向差异的实例。从顶级经济学和金融期刊中提取10,490个因果三元组(治疗-结果对,具有实证验证的效果方向),识别出1,056个意识形态冲突实例,并对20种最先进的LLM进行评估,以预测实证支持的因果方向。研究发现,意识形态冲突实例比非冲突实例更难处理,且在18种模型中,当实证验证的因果符号与干预导向预期一致时,准确性更高。此外,当模型出错时,其错误预测倾向于干预导向,且这种方向性偏差无法通过一次性的上下文提示消除。这些结果表明,LLM在意识形态冲突的经济问题上不仅准确性较低,而且在某一意识形态方向上系统性地不可靠,凸显了在高风险经济和政策设置中需要方向感知评估的必要性。

英文摘要

Do large language models (LLMs) exhibit systematic ideological bias when reasoning about economic causal effects? As LLMs are increasingly used in policy analysis and economic reporting, where directionally correct causal judgments are essential, this question has direct practical stakes. We present a systematic evaluation by extending the EconCausal benchmark with ideology-contested cases - instances where intervention-oriented (pro-government) and market-oriented (pro-market) perspectives predict divergent causal signs. From 10,490 causal triplets (treatment-outcome pairs with empirically verified effect directions) derived from top-tier economics and finance journals, we identify 1,056 ideology-contested instances and evaluate 20 state-of-the-art LLMs on their ability to predict empirically supported causal directions. We find that ideology-contested items are consistently harder than non-contested ones, and that across 18 of 20 models, accuracy is systematically higher when the empirically verified causal sign aligns with intervention-oriented expectations than with market-oriented ones. Moreover, when models err, their incorrect predictions disproportionately lean intervention-oriented, and this directional skew is not eliminated by one-shot in-context prompting. These results highlight that LLMs are not only less accurate on ideologically contested economic questions, but systematically less reliable in one ideological direction than the other, underscoring the need for direction-aware evaluation in high-stakes economic and policy settings.

URL PDF HTML 收藏
2602.23242 2026-07-14 cs.AI 版本更新

A Model-Free Universal AI

无模型通用人工智能

Yegon Kim, Juho Lee

机构 * Graduate School of AI, KAIST(韩国科学技术院人工智能研究生院)

AI总结 提出首个在通用强化学习中证明渐近ε最优的无模型智能体AIQI,通过分布动作值函数的通用归纳实现,并扩展了Self-AIXI的渐近最优性证明。

Comments 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)

详情
AI中文摘要

在通用强化学习中,所有已建立的最优智能体,包括AIXI,都是基于模型的,显式维护和使用环境模型。本文介绍了具有Q归纳的通用人工智能(AIQI),这是首个被证明在通用RL中渐近ε最优的无模型智能体。AIQI对分布动作值函数进行通用归纳,而不是像先前工作那样对策略或环境进行归纳。在“真理颗粒”条件下,我们证明了AIQI是强渐近ε最优和渐近ε贝叶斯最优的。我们还应用我们的新颖证明技术,在没有特别假设的情况下证明了Self-AIXI的渐近ε最优性。我们的结果显著扩展了已知通用智能体的多样性。

英文摘要

In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using environment models. This paper introduces Universal AI with Q-Induction (AIQI), the first model-free agent proven to be asymptotically $\varepsilon$-optimal in general RL. AIQI performs universal induction over distributional action-value functions, instead of policies or environments like previous works. Under a grain of truth condition, we prove that AIQI is strong asymptotically $\varepsilon$-optimal and asymptotically $\varepsilon$-Bayes-optimal. We also apply our novel proof techniques to show asymptotic $\varepsilon$-optimality of Self-AIXI without any ad-hoc assumptions. Our results significantly expand the diversity of known universal agents.

URL PDF HTML 收藏
2601.13132 2026-07-14 cs.CV 版本更新

SplatReasoner: Enhancing Embodied Reasoning and Grounding by Novel View Synthesis

SplatReasoner:通过新颖视图合成增强具身推理与基础能力

Kim Yu-Ji, Dahye Lee, Kim Jun-Seong, Nam Hyeon-Woo, GeonU Kim, Yongjin Kwon, Yu-Chiang Frank Wang, Jaesung Choe, Tae-Hyun Oh

机构 * POSTECH KAIST(韩国科学技术院) ETRI(韩国电子电信研究院) NVIDIA(英伟达)

AI总结 研究针对视觉语言模型应用于具身场景理解受固定视角限制的问题,提出SplatReasoner框架,利用3D高斯点云将新颖视图合成引入推理过程,经实验验证该方法能提升具身推理和3D基础能力。

Comments Accepted at ECCV 2026. Project page: https://splatreasoner.github.io/

详情
AI中文摘要

视觉语言模型(VLMs)在图像和视频上展现出强大推理能力,但应用于具身场景理解时,常受限于情景RGB-D记忆中的固定视角。这些观察可能因遮挡、物体截断、视野受限或视图组合不佳而无法捕捉与查询相关的证据。我们提出SplatReasoner框架,通过利用3D高斯点云(3DGS)将新颖视图合成引入VLM推理过程。给定关于3D场景的用户查询,SplatReasoner检索相关观察并合成查询条件视角,以揭示回答查询和在3D中定位所指实体所需的视觉证据。实验表明,查询条件新颖视图合成在固定视角记忆和语言嵌入3DGS基线之上,提升了具身推理和3D基础能力。

英文摘要

Vision-Language Models (VLMs) have demonstrated strong reasoning capabilities over images and videos, yet their application to embodied scene understanding often constrained by the fixed viewpoints stored in episodic RGB-D memories. These observations may fail to capture query-relevant evidence due to occlusions, object truncation, restricted fields of view, or suboptimal view composition. We present SplatReasoner, a framework that introduces novel view synthesis into the VLM reasoning process by leveraging 3D Gaussian Splatting (3DGS). Given a user query about a 3D scene, SplatReasoner retrieves relevant observations and synthesizes query-conditioned viewpoints that reveal the visual evidence needed to answer the query and ground the referred entities in 3D. Experiments show that query-conditioned novel view synthesis improves both embodied reasoning and 3D grounding over fixed-view memory and language-embedded 3DGS baselines.

URL PDF HTML 收藏
2607.09362 2026-07-13 cs.CV cs.AI 新提交

CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation

CtrlVTON:通过视觉实例提示分割实现可控虚拟试穿

Seungyong Lee, Hyun Jun Jang, Sangoh Kim, Sungjoon Park

机构 * NXN Labs(NXN实验室) KAIST(韩国科学技术院)

AI总结 研究旨在解决虚拟试穿中用户对服装穿着方式控制不足的问题。提出通过VIP - SAM解决视觉实例提示分割,引入CtrlVTON可控框架,将试穿转为图像编辑问题并添加分割掩码控制布局。二者在各自任务达先进水平,CtrlVTON能更忠实地遵循用户布局且保证服装逼真度。

Comments 13 + 17 pages, 20 figures

详情
AI中文摘要

虚拟试穿(VTO)在将服装逼真地转移到目标人物身上方面取得了重大进展。然而,大多数系统让用户几乎无法控制服装的穿着方式,包括尺寸(宽松或合身)、款式(如塞进或不塞进、敞开或闭合)以及在身体上的空间位置。我们通过两个互补的贡献来解决这一差距。首先,我们通过VIP - SAM定义并解决视觉实例提示分割:给定一件服装的平铺图像,在穿着该服装的人的照片中分割出特定实例。这是一个实例级任务,不同于通常研究的类别级分割。其次,我们引入CtrlVTON,一个可控的VTO框架,将试穿重新定义为图像编辑问题,并添加分割掩码作为对服装布局的像素级控制,包括款式、尺寸和在身体上的空间位置。VIP - SAM和CtrlVTON在各自任务上均取得了领先成果。特别是,CtrlVTON生成的图像比最强的专有编辑系统更忠实地遵循用户提供的布局,同时在服装逼真度上与之匹配。

英文摘要

Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g., tucked in or untucked, open or closed), and spatial placement on the body. We address this gap with two complementary contributions. First, we define and solve Visual-Instance-Prompt Segmentation via VIP-SAM: given a flatlay image of a garment, segment that specific instance in a photograph of a person wearing it. This is an instance-level task, distinct from the typically studied category-level segmentation. Second, we introduce CtrlVTON, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body. VIP-SAM and CtrlVTON each achieve state-of-the-art results on their respective tasks. In particular, CtrlVTON generates images that follow user-provided layouts far more faithfully than the strongest proprietary editing systems while matching them on garment fidelity.

URL PDF HTML 收藏
2607.09263 2026-07-13 cs.CV 新提交

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval

语义难度并非视觉难度:用于手语检索的符号感知硬负样本挖掘

Junmyeong Lee, Chan Hur, ChangSu Choi, Sukmin Cho, Fitsum Gaim, Eui Jun Hwang, Hoyun Song, KyungTae Lim

机构 * School of Computing(计算机学院) Graduate School of Culture Technology(文化技术研究生院) Korea Advanced Institute of Science and Technology(韩国科学技术院) ETRI Medical Informatics Laboratory(电子通信研究院医学信息学实验室)

AI总结 研究手语检索在细粒度场景的问题,提出符号感知硬负样本挖掘方法,通过在嵌入空间基于视觉易混淆性构建硬负样本,实验证明该方法能提升细粒度检索性能且保持粗粒度准确性。

Comments Accepted to ACL 2026 main

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026, pages 28262-28277

详情
AI中文摘要

手语检索(SLRet)能高效访问手语内容,但在需区分视觉相似符号的细粒度场景中仍很脆弱。我们表明此限制并非源于模型能力,而是无效的硬负样本监督。具体而言,我们将细粒度检索失败表述为负分布不匹配:语义不同但视觉上易混淆的符号很少被视为硬负样本,而现有基于文本的挖掘策略无法捕捉这种视觉模糊性。为解决此问题,我们提出符号感知硬负样本挖掘(SAN),它基于手语嵌入空间中的视觉易混淆性构建硬负样本。在PHOENIX - 2014T上的实验表明,SAN在保持粗粒度准确性的同时显著提高了细粒度检索性能,突出了在手语检索中使负样本监督与视觉模糊性对齐的重要性。

英文摘要

Sign Language Retrieval (SLRet) enables efficient access to sign language content but remains fragile in fine-grained scenarios where visually similar signs must be distinguished. We show that this limitation does not stem from model capacity, but from ineffective hard negative supervision. Specifically, we formulate fine-grained retrieval failures as a negative distribution mismatch: semantically distinct yet visually confusable signs are rarely treated as hard negatives, while existing text-based mining strategies fail to capture such visual ambiguity. To address this issue, we propose Sign-Aware Hard Negative Mining (SAN), which constructs hard negatives based on visual confusability in the sign embedding space rather than linguistic similarity. Experiments on PHOENIX-2014T demonstrate that SAN substantially improves fine-grained retrieval performance while preserving coarse-grained accuracy, highlighting the importance of aligning negative supervision with visual ambiguity in sign language retrieval.

URL PDF HTML 收藏
2607.09167 2026-07-13 cs.LG math.OC 新提交

Understanding Schedule-Free Methods in Nonconvex Optimization: Rate Guarantees and Escaping Saddles

理解非凸优化中无调度方法:速率保证与逃离鞍点

Jiseok Chae, Donghwan Kim

机构 * KAIST(韩国科学技术院) Natural Science Research Institute(自然科学研究所) Department of Mathematical Sciences(数学科学系)

AI总结 研究非凸优化中无调度方法,通过李雅普诺夫分析给出最坏情况收敛速率分析,还将无调度梯度下降建模为非自治动力系统,证明其在微小扰动下可避免严格鞍点,解释了该方法性能良好的原因。

Comments 44+7 pages, 2 figures

详情
AI中文摘要

无调度方法因减轻学习率调度器设计和调整负担而备受关注,其性能有时优于有调度的优化器。尽管实证结果良好,但非凸优化中的收敛理论仍未充分探索。本文对标准形式的无调度梯度下降和无调度随机梯度下降进行最坏情况分析,基于李雅普诺夫分析表明它们达到一阶方法的最优最坏情况收敛速率,还证明了无调度梯度下降在微小一次性扰动下可避免严格鞍点,有助于更好理解其性能。

英文摘要

Schedule-Free methods have attracted growing interest for alleviating the burden of designing and tuning a learning rate scheduler, while matching and sometimes even outperforming optimizers with tuned schedulers. Despite their strong empirical results, their convergence theory in nonconvex optimization, where modern machine learning objectives typically arise, has remained largely unexplored. In this paper, we provide worst-case analyses of Schedule-Free gradient descent and Schedule-Free stochastic gradient descent, in their standard form and without auxiliary modifications or restrictive conditions, for smooth but possibly nonconvex objectives. Based on a Lyapunov analysis derived from the continuous-time limiting ordinary differential equation associated with these methods, we show that Schedule-Free gradient descent and Schedule-Free stochastic gradient descent achieve the optimal worst-case convergence rates attainable among first-order methods. We further formulate Schedule-Free gradient descent as a nonautonomous dynamical system and prove strict-saddle avoidance under an arbitrarily small one-time perturbation. These theoretical results provide a better understanding of the strong performance that Schedule-Free methods demonstrate.

URL PDF HTML 收藏