arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

The Chinese University of Hong Kong(香港中文大学)

至 收录 2417
2607.17766 2026-07-21 cs.CL 新提交

When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation

何时使用额外上下文:基于证据的术语适应在同步语音翻译中的应用

Zeyu Yang, Satoshi Nakamura

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

AI总结 研究同步语音翻译中额外上下文的使用,提出基于证据的术语适应框架EGTA,通过构建术语记忆、选择候选术语来调整决策空间,无需全模型微调,在多个指标上有提升,改进与特定论文证据对齐有关。

详情
AI中文摘要

额外上下文对技术讲座的同步语音翻译很有价值,但将整个文档上下文注入每个流片段往往过于粗糙。通过诊断实验发现上下文增益主要来自特定论文术语恢复而非统一语义增强。因此提出EGTA框架,它构建文档术语记忆,根据当前流状态选择紧凑候选术语,并仅使用所选术语调整ASR/语音端和解码器端决策空间。EGTA可在多种同步语音翻译设置中实例化,无需全模型微调。在相关评估套件上评估,结果显示其在多个指标上有提升,且改进与特定论文证据对齐相关。

英文摘要

Extra context is valuable for simultaneous speech translation of technical talks, but injecting the entire document context into every streaming segment is often too coarse. Through diagnostic experiments, we find that context gains mainly come from paper-specific terminology recovery rather than uniform semantic enhancement. We therefore propose EGTA, an Evidence-Grounded Terminology Adaptation framework that builds a document terminology memory, selects compact candidate terms conditioned on the current streaming state, and adapts ASR/speech-side and decoder-side decision spaces using only the selected terms. EGTA can be instantiated in cascaded, end-to-end, and generation-only SimulST settings without full-model fine-tuning. We evaluate EGTA on an ACL technical-talk SimulST evaluation suite consisting of MCIF-dev and ACL60/60-dev. On MCIF-dev, EGTA-RG improves BLEU by +1.05/+0.59, XCOMET-XL by +0.019/+0.006, named-entity recall by +79\%/+73\% relative, and acronym recall by +0.099/+0.171 on En$\rightarrow$Zh and En$\rightarrow$De. Across MCIF-dev latency settings, EGTA consistently improves XCOMET-XL, named-entity recall, and acronym recall. External validation on ACL60/60-dev further shows consistent terminology-recall gains without additional fine-tuning. Shuffled-memory controls and activation audits provide evidence that the improvements are tied to paper-specific evidence alignment rather than generic context prompting.

URL PDF HTML 收藏
2607.17095 2026-07-21 cs.AI 新提交

Fourier Geometric Wind Power Forecasting with Numerical Weather Prediction

基于数值天气预报的傅里叶几何风力发电预测

Shiyuan Piao, Fan Zehui, Yang Liu, Hong Cheng, Juepeng Zheng, Jie Zhou, Fugee Tsung

机构 * The Hong Kong University of Science and Technology(香港科技大学) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Goldwind Science and Technology Co.,Ltd(金风科技股份有限公司)

AI总结 研究针对风力发电预测难题,提出多模态框架,整合SCADA数据与NWP预测。先分解输入特征,再用几何编码器和傅里叶神经算子建模,实验表明该模型优于现有基线,凸显基于物理设计的有效性。

详情
AI中文摘要

准确的短期风力发电预测对电网稳定性和运营规划至关重要,但由于大气条件与涡轮机动力学之间的复杂相互作用,这一任务仍具有挑战性。现有方法未能有效整合天气预报与风力涡轮机数据(即SCADA),导致解决方案欠佳。为解决此问题,我们引入了一个多模态框架,将基于历史点的SCADA数据与基于网格的数值天气预报(NWP)预测相结合,这因异构输入和复杂的物理风力涡轮机相互作用而颇具挑战。我们的方法首先将输入明确分解为标量和矢量特征,以更好地捕捉特定地点和几何相关性,然后采用几何编码器从风矢量中提取旋转不变特征。我们还利用了傅里叶神经算子(FNO)架构,它在频域中执行全局卷积,以有效建模远程时空关系。在三个实际风电场进行的广泛实验表明,我们的模型始终优于现有基线,凸显了其基于物理的设计的有效性。我们方法的核心实现可在指定网址公开获取。

英文摘要

Accurate short-term wind power forecasting is essential for grid stability and operational planning, yet remains challenging due to the complex interactions between atmospheric conditions and turbine dynamics. However, existing methods fail to effectively incorporate weather forecasting with wind turbine data (i.e., SCADA), leading to suboptimal solutions. To address this, we introduce a multimodal framework that integrates historical point-based SCADA data with grid-based Numerical Weather Prediction (NWP) forecasts, which is challenging due to heterogeneous input and the complex physical wind-turbine interactions. Our approach first explicitly decomposes inputs into scalar and vector features to better capture both site-specific and geometric dependencies and then incorporates a geometric encoder to extract rotation-invariant features from wind vectors. We further leverages a Fourier Neural Operator (FNO) architecture, which performs global convolutions in the frequency domain to efficiently model long-range spatiotemporal relationships. Extensive experiments on three real-world wind farms, with weather forecasting data, demonstrate that our model consistently outperforms state-of-the-art baselines, highlighting the effectiveness of its physically-informed design. The core implementation of our method is publicly available at: https://github.com/shawn-sypiao/GWPF.

URL PDF HTML 收藏
2607.16859 2026-07-21 cs.CV 新提交

Dataset Distillation by Influence Matching

通过影响匹配进行数据集蒸馏

Haoru Tan, Wang Wang, Sitong Wu, Xiuzhe Wu, Yangtian Sun, Chirui Chang, Shaofeng Zhang, Xiaojuan Qi

机构 * HKU(香港大学) CUHK(香港中文大学) Stanford(斯坦福大学)

AI总结 从结果角度重审数据集蒸馏,提出影响匹配方法,通过可微估计器量化参数变化,学习合成集使影响与真实数据集匹配,在分类和视觉语言蒸馏任务中表现出色,超越基线。

Journal ref CVPR 2026

详情
AI中文摘要

我们从以结果为中心的角度重新审视数据集蒸馏。影响匹配(Inf-Match)不是对齐过程代理(逐步骤梯度或训练轨迹),而是对齐训练的最终结果:它学习一个紧凑的合成集,其对收敛参数的影响与完整数据集的影响相匹配。具体而言,我们引入了一个完全可微的样本级影响估计器,通过展开优化动态并应用一阶泰勒近似,在线性时间内量化添加或删除数据导致的参数变化。然后通过最小化合成集与真实数据集影响之间的不匹配来学习合成集,实现结果对齐而非启发式过程模仿。Inf-Match在标准分类基准上实现了最佳准确率,在图像/文本检索任务上也优于强过程匹配基线。代码将通过此https URL发布。

英文摘要

We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (per-step gradients or training trajectories), Influence Matching (Inf-Match) aligns the final outcome of training: it learns a compact synthetic set whose effect on the converged parameters matches that of the full dataset. Concretely, we introduce a fully differentiable, sample-level influence estimator that quantifies parameter shifts from adding or removing data, without time-consuming inverse-Hessian products or convexity assumptions. The estimator runs in linear time by unrolling the optimization dynamics and applying a first-order Taylor approximation. We then learn the synthetic set by minimizing the mismatch between its influence and that of the real dataset, yielding outcome alignment rather than heuristic process imitation. Inf-Match delivers the best accuracy across standard classification benchmarks. For instance, on Tiny-ImageNet (IPC=10), Inf-Match attains 31.5\%, a +4.7\% improvement over NCFM. Beyond classification, Inf-Match scales to vision-language distillation on Flickr30K, outperforming strong process-matching baselines. For instance, with 200 to 1000 synthetic samples, our method achieved a leading impressive average on image/text retrieval tasks, higher than NCFM by 2.5\%. The code will be released via https://github.com/hrtan/infmatch.

URL PDF HTML 收藏
2607.16657 2026-07-21 cs.SD 新提交

HARP: Harmonic-Aware Residual Partitioning for Neural Audio Codecs

HARP:用于神经音频编解码器的谐波感知残差划分

Qiaoyu Yang, Lixing He, Binyue Deng, Weifeng Zhao

机构 * Georgia Institute of Technology(佐治亚理工学院) The Chinese University of Hong Kong(香港中文大学) Tencent Music Entertainment(腾讯音乐娱乐集团)

AI总结 研究针对神经音频编解码器中码本频谱纠缠等问题,提出HARP训练策略,将RVQ阶段分组,在解码器能访问低频时各小组细化目标频带,重建泛音保留连贯性,该策略无需架构改变,性能优于标准RVQ和并行分解。

Comments Accepted to Interspeech 2026

详情
AI中文摘要

具有残差向量量化(RVQ)的神经音频编解码器通常对所有频率一视同仁,导致其码本频谱纠缠。截断阶段会去除不可预测的频率混合。并行频带分解通过将音频拆分为独立频带来解决此问题,但会使潜在空间碎片化并失去跨频率连贯性。我们引入了HARP(谐波感知残差划分),这是一种训练策略,将RVQ阶段划分为按频率排序的组,每个组在解码器仍可访问所有低频的同时细化其目标频带。泛音在基音的背景下重建,保留了并行方法所失去的连贯性。HARP无需架构更改,仅修改训练损失,推理与标准RVQ相同。在语音、音乐和一般音频上,HARP优于标准RVQ和并行分解。MUSHRA听力测试也显示出感知上的改进。

英文摘要

Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio into independent bands, but fragments the latent space and loses cross-frequency coherence. We introduce HARP (Harmonic-Aware Residual Partitioning), a training strategy that partitions RVQ stages into frequency-ordered groups where each group refines its target band while the decoder retains access to all lower frequencies. Overtones are reconstructed in the context of their fundamentals, preserving coherence that parallel methods lose. HARP requires no architectural changes; it only modifies the training loss, leaving inference identical to standard RVQ. On speech, music, and general audio, HARP outperforms both standard RVQ and parallel decomposition. MUSHRA listening tests also show perceptual improvements.

URL PDF HTML 收藏
2607.16599 2026-07-21 cs.SD cs.MM eess.AS 新提交

Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves

一个分数够吗?用时间分数曲线评估歌曲的演唱质量

Yishan Lv, Jing Luo, Xinyu Yang, Zhizheng Wu

机构 * Xi’an Jiaotong University(西安交通大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

AI总结 研究针对全长歌曲演唱质量评估难题,提出SongSQA两阶段框架。第一阶段用伪标签训练段分数预测器,第二阶段聚合器整合特征与分数生成嵌入并捕捉联系,有效提升评估效果,在KTAU上相对提高13.95%。

详情
AI中文摘要

演唱质量评估(SQA)在实际多媒体应用和音乐人工智能系统中变得越来越重要,但现有研究主要集中在短演唱片段,对全长歌曲的评估还不够。全长歌曲SQA需要对演唱质量在不同音频段的变化以及这些局部变化如何影响整体演唱表现评估进行建模。此外,段级注释的稀缺使得有效监督具有挑战性。为应对这些挑战,我们提出了SongSQA,这是一个用于全长歌曲SQA的两阶段框架。第一阶段,使用预训练教师模型生成的伪标签训练段分数预测器,无需手动段注释即可进行段级演唱质量预测。第二阶段,歌曲质量聚合器将段特征和预测的段分数集成到统一的段嵌入中,并使用可学习的歌曲嵌入和自注意力来捕捉段级演唱表现与整体歌曲质量之间的联系。实验结果证明了SongSQA对全长歌曲SQA的有效性,在KTAU上比最强基线相对提高了13.95%,同时在所有数据集上持续改进其他评估指标。

英文摘要

Singing Quality Assessment (SQA) has become increasingly important for practical multimedia applications and Music AI systems, yet existing studies predominantly focus on short singing clips and remain insufficient for full-length songs. Unlike clip-level assessment, full-length song SQA requires modeling how singing quality varies across different audio segments and how these local variations influence the overall evaluation of vocal performance. Moreover, the scarcity of segment-level annotations makes effective supervision challenging, as directly assigning a single overall score label to every segment tends to treat different segment qualities as equivalent. To address these challenges, we propose SongSQA, a two-stage framework for full-length song SQA. In the first stage, a Segment Score Predictor is trained with pseudo labels generated by a pre-trained teacher model, enabling segment-level singing quality prediction without requiring manual segment annotations. In the second stage, a Song Quality Aggregator integrates segment features and predicted segment scores into unified segment embeddings, and employs a learnable song embedding together with self-attention to capture the connection between segment-level vocal performance and overall song quality. In this way, SongSQA dynamically aggregates critical quality cues across the song to produce a holistic quality prediction, while also generating a temporal segment-level quality curve. Experimental results demonstrate the effectiveness of SongSQA for full-length song SQA, achieving up to a 13.95% relative improvement in KTAU over the strongest baseline, while consistently improving other evaluation metrics across all datasets.

URL PDF HTML 收藏
2607.16577 2026-07-21 cs.CV cs.GR 新提交

CNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape Representation

CNS-Edit++:基于耦合神经形状表示的类别无关3D编辑

Jingyu Hu, Weilong Yan, Zhengzhe Liu, Haipeng Li, Ka-Hei Hui, Hao, Zhang, Chi-Wing Fu

机构 * The Chinese University of Hong Kong(香港中文大学) Lingnan University(岭南大学) National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology(香港科技大学) Autodesk AI Lab(欧特克人工智能实验室) Simon Fraser University(西蒙弗雷泽大学)

AI总结 研究提出基于耦合神经形状表示和神经特征体积优化的潜在空间3D形状编辑框架CNS-Edit++,能在特定类别和类别无关模型上实例化,有多种编辑操作符及区域控制机制,经评估其性能优于现有方法。

详情
AI中文摘要

本文提出了一个基于耦合神经形状(CNS)表示和神经特征体积优化的潜在空间3D形状编辑框架。该工作将基于耦合神经形状优化的CNS-Edit扩展到CNS-Edit++,通过将特定类别的耦合表示推广到使用基础模型的类别无关3D形状编辑。耦合神经形状(CNS)表示将捕获高级形状语义的全局潜在代码与为局部形状操作提供空间上下文的3D神经特征体积耦合。然后制定了一个耦合神经形状优化过程,以根据给定的编辑操作共同优化这两个组件。该框架可以在特定类别的3D反演模型和类别无关的3D基础模型上实例化。提供了各种形状编辑操作符,并引入两种互补的区域控制机制以保留编辑区域外的区域。不同3D生成模型的广泛定量和定性评估证明了该方法优于现有解决方案的强大能力。

英文摘要

This paper presents a latent-space 3D shape editing framework built upon a coupled neural shape (CNS) representation and a neural feature volume optimization. This work extends CNS-Edit, built on Coupled Neural Shape optimization, to CNS-Edit++, by generalizing the category-specific coupled representation to category-agnostic 3D shape editing with foundation models. The Coupled Neural Shape (CNS) representation couples a global latent code that captures high-level shape semantics with a 3D neural feature volume that provides spatial context for local shape manipulation. Then we formulate a coupled neural shape optimization procedure that co-optimizes these two components subject to a given editing operation. Our framework can be instantiated on both the category-specific 3D inversion model and category-agnostic 3D foundation models. We provide various shape editing operators, including copy, resize, delete, mix, point-wise drag, and region-wise drag, each of which is formulated as an objective to guide the CNS optimization. To preserve regions outside the editing area, we further introduce two complementary region-wise control mechanisms, i.e., KV-cache replacement and latent feature regularization. Extensive quantitative and qualitative evaluations across different 3D generative models demonstrate the strong capabilities of our approach over state-of-the-art solutions.

URL PDF HTML 收藏
2607.16427 2026-07-21 cs.CL 新提交

Multi-level context Modeling for consistent expert selection in Mixture-of-Experts

用于混合专家模型中一致专家选择的多层次上下文建模

Shuhan Huang, Naifan Zhang, Yuanbo Tang, Yang Li, Wai Kin Victor Chan

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) School of AI, The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)人工智能学院)

AI总结 研究混合专家模型中专家选择问题,提出多层次上下文融合MoE框架,通过整合跨层语义聚合和局部令牌级交互信号构建上下文感知表示,提升路由一致性和下游性能。

详情
AI中文摘要

混合专家模型(MoE)通过将令牌路由到一小部分专家来实现Transformer模型的高效扩展。然而,现有路由器通常基于浅层或孤立的令牌表示来进行专家选择,这往往会在各层产生不稳定且语义不一致的路由决策。在这项工作中,我们从表示角度重新审视专家选择,并将上下文不完整性识别为限制有效专家专业化的关键瓶颈。为解决此问题,我们提出了多层次上下文融合MoE(MCF-MOE)框架,该框架通过整合来自跨层语义聚合和局部令牌级交互的互补信号来构建上下文感知表示,从而实现更具信息性和一致性的专家选择。在语言建模和理解基准测试上的实验表明,MCF-MOE相对于强大的MoE基线持续提高了路由一致性和下游性能,突出了上下文完整性在专家路由中的重要性。代码可在该https URL获取。

英文摘要

Mixture-of-Experts (MoE) enables efficient scaling of Transformer models by routing tokens to a small subset of experts. However, existing routers typically condition expert selection on shallow or isolated token representations, which often produce unstable and semantically inconsistent routing decisions across layers. In this work, we revisit expert selection from a representation perspective and identify context incompleteness as a key bottleneck limiting effective expert specialization. To address this issue, we propose Multi-level Context Fusion MOE (MCF-MOE), a framework that constructs context-aware representations by integrating complementary signals from cross-layer semantic aggregation and local token-level interactions, enabling more informative and consistent expert selection. Experiments on language modeling and understanding benchmarks demonstrate that MCF-MOE consistently improves routing consistency and downstream performance over strong MoE baselines, highlighting the importance of contextual completeness in expert routing. The code is available at https://anonymous.4open.science/r/MCFMOE.

URL PDF HTML 收藏
2607.16401 2026-07-21 cs.CV 新提交

Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

Apple-$π$:基于视频对基于法律的物理智能进行思维基准测试

Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang, Shulin Tian, Ziang Cao, Weichen Fan, Ziqi Huang, Yuhao Dong, Hao Li, Zhaoxi Chen, Zhongang Cai, Lei Yang, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S-Lab) The Chinese University of Hong Kong(香港中文大学)

AI总结 研究针对视频模型评估仅看输出层面的不足,提出Apple-PI基准,含数据集、协议和评估套件三部分,通过对11个模型测试发现其距可靠世界模拟器差距大,并揭示了一些瓶颈与差距,为指导视频模型发展提供诊断基础。

详情
AI中文摘要

现代视频生成模型被誉为对物理定律有内在理解的新兴世界模型。然而现有基准大多仅在输出层面评估物理合理性,未验证模型是否通过忠实的、基于定律的推理过程得出结果。我们引入了Apple-PI,首个明确以物理定律为锚点的视频模型评估基准。它由三部分组成:包含400个视频的Orchard数据集,涵盖经典力学十个规范任务,区分单定律与多定律任务;基于科学推理的三阶段基准协议,包括感知、公式化和推导,对首帧进行信息图表注释并使用帧链提示;结合基于MLLM的主观评分与基于物理定律的客观度量的混合评估套件。对11个模型的基准测试表明当前视频模型距可靠的基于定律的世界模拟器仍有很大差距,我们的分析还揭示了感知到公式化再到推导的瓶颈、多定律状态转移薄弱以及模拟到现实的持续差距。这些发现使Apple-PI成为指导未来视频模型成为具有基于法律的物理智能的世界模型的诊断基础。

英文摘要

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-PI comprises three components. 1) Orchard: a dataset of 400 videos covering ten canonical tasks in classical mechanics. It separates single-law tasks for confounder-free diagnosis from multi-law tasks for probing generalization. 2) Benchmark Protocol: a three-stage protocol based on scientific reasoning, including Perception, Formulation, and Deduction. It uses chain-of-frames prompting on infographic-annotated first frames, treating the generated video as the model's visible reasoning trace. 3) Evaluation Suite: a hybrid evaluation suite that combines MLLM-based subjective scoring with physics-law-grounded objective measures. This enables stage-resolved diagnosis of not only whether a model fails, but where it fails. Benchmarking 11 models shows that current video models remain far from reliable law-grounded world simulators, with the best video model scoring only 0.473. Our stage-, pillar-, and source-resolved analyses further expose a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. These findings position Apple-PI as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.

URL PDF HTML 收藏
2607.16295 2026-07-21 cs.CV cs.AI 新提交

Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm

基于群对比前向算法的涌现分层单语义神经元

Yiming Tang, Qinglin Qi, Zhaoqian Yao, Harshvardhan Saini, Dianbo Liu

机构 * National University of Singapore(新加坡国立大学) Lund University(隆德大学) Chinese University of Hong Kong(香港中文大学) Indian Institute of Technology, Dhanbad(印度理工学院(丹巴德分校))

AI总结 研究探讨神经网络表示的可解释性,针对稀疏字典学习范式的局限,提出群对比前向算法GCFF,通过架构约束实现单语义性,能捕捉非线性概念,在CLIP表示上表现良好,还能从头训练网络并在图像分类基准中达最优性能。

详情
AI中文摘要

机械可解释性在理解神经网络表示方面取得了显著进展,稀疏字典学习(SDL)方法是核心范式,但存在局限性。我们假设存在不同的单语义性途径,生物视觉系统有高度选择性的神经元分层组织,源于局部、逐层学习规则。为此提出群对比前向算法(GCFF),通过架构约束而非稀疏性实现单语义性,能捕捉非线性概念。在CLIP表示上,单个训练的GCFF模块可恢复抽象度随深度递增的单语义神经元,且无需稀疏约束或抽象级别监督。此外,GCFF能从头训练网络,在各种图像分类基准上达到前向算法的最优性能。

英文摘要

Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning (SDL) methods, most prominently sparse autoencoders, as a central paradigm. However, recent work has reported several limitations of this paradigm: SDL objectives are non-identifiable; SDL methods rely heavily on the Linear Representation Hypothesis; and a growing body of evidence points to concepts that are encoded non-linearly and are therefore not expressible as any single direction. We hypothesise that a different route to monosemanticity is available. Biological visual systems exhibit highly selective neurons organised into hierarchies of increasing abstraction, and this organisation emerges from local, layer-wise learning rules rather than from a global error signal; we therefore ask whether a biologically plausible learning algorithm will likewise yield monosemantic neurons. To test this, we propose Group-Contrastive Forward-Forward (GCFF), a forward-forward training algorithm that combines class-specific routing with within-class contrastive objectives, reaching monosemanticity through architectural constraints rather than sparsity. Because GCFF attaches multiple non-linear layers to the representation under study, its neurons can therefore capture the non-linear concepts. On CLIP representations, a single trained GCFF module recovers monosemantic neurons whose abstraction increases progressively with depth, reaching environmental properties that hold independently of an image's foreground, without any sparsity constraint or supervision of abstraction level. We further demonstrate that GCFF can train networks from scratch, achieving state-of-the-art performance among forward-forward algorithms on various image classification benchmarks.

URL PDF HTML 收藏
2606.30362 2026-07-21 cs.RO cs.AI cs.CV 版本更新

ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control

ReactiveBFM:面向通用人形全身控制的反应式闭环运动规划

Xiao Chen, Weishuai Zeng, Xiaojie Niu, Zirui Wang, Jianan Li, Huayi Wang, Furui Xu, Jiahe Chen, Weixiang Zhong, Lihe Ding, Kailin Li, Jiangmiao Pang, Tai Wang, Tianfan Xue, Jingbo Wang

机构 * The Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室)

AI总结 提出ReactiveBFM,一种实时闭环规划控制框架,通过调度前缀采样课程和异步重规划机制,解决生成式运动规划器与高频跟踪之间的延迟不匹配和暴露偏差问题,实现人形机器人全身协调和零样本反应式运动。

Comments Project page: https://xiao-chen.tech/reactivebfm/

详情
AI中文摘要

虽然当前的行为基础模型(BFMs)为人形机器人提供了鲁棒的控制先验,但它们仅执行预定义的参考运动。因此,它们容易受到环境变化的影响,并且无法实现反应式全身协调。将它们与生成式运动规划器简单级联无法实现真正的反应性,因为不可避免的跟踪差异会导致致命的累积暴露偏差。为了弥合这一差距,我们提出了ReactiveBFM,一个实时闭环规划控制框架。其核心是通过调度前缀采样课程有效缓解暴露偏差,迫使生成式规划器从非完美的物理状态而非真实轨迹中主动学习错误恢复行为。系统性地,为了解决自回归规划与高频跟踪之间的严重延迟不匹配,我们引入了一种异步重规划机制。结合轨迹分块以时间上集成空间参考,我们的系统保证了无物理抖动的时空流畅执行。在Unitree G1人形机器人上部署后,ReactiveBFM在大量文本条件的闭环运动中展示了前所未有的物理敏捷性。值得注意的是,ReactiveBFM实现了零样本移动目标到达,展示了复杂的全身协调和即时重规划。在严重扰动下的仿真到仿真基准测试中,ReactiveBFM达到了93.1%的成功率,显著优于级联开环基线28.6%。

英文摘要

While current Behavior Foundation Models (BFMs) provide robust control priors for humanoids, they only execute pre-defined reference motions. As a result, they are vulnerable to environmental shifts and incapable of reactive whole-body coordination. Naively cascading them with generative motion planners fails to achieve true reactivity, as inevitable tracking discrepancies induce fatal cumulative exposure bias. To bridge this gap, we propose ReactiveBFM, a real-time closed-loop planning-control framework. At its core, we effectively mitigate exposure bias via a scheduled prefix sampling curriculum, forcing the generative planner to actively learn error-recovery behaviors from imperfect physical states rather than ground-truth trajectories. Systematically, to reconcile the severe latency mismatch between auto-regressive planning and high-frequency tracking, we introduce an asynchronous replanning mechanism. Combined with trajectory chunking to temporally ensemble spatial references, our system guarantees spatio-temporally fluid execution without physical jitter. Deployed on the Unitree G1 humanoid, ReactiveBFM demonstrates unprecedented physical agility across a vast repertoire of text-conditioned closed-loop motions. Notably, ReactiveBFM achieves zero-shot moving target reaching, showcasing intricate whole-body coordination and on-the-fly replanning. In sim-to-sim benchmarking under severe perturbations, ReactiveBFM achieves a 93.1% success rate, significantly outperforming cascaded open-loop baselines by 28.6%.

URL PDF HTML 收藏
2606.19341 2026-07-21 cs.CV cs.CL cs.SD 版本更新

Native Active Perception as Reasoning for Omni-Modal Understanding

原生主动感知作为全模态理解的推理

Zhenghao Xing, Ruiyang Xu, Yuxuan Wang, Jinzheng He, Ziyang Ma, Qize Yang, Yunfei Chu, Jin Xu, Junyang Lin, Chi-Wing Fu, Pheng-Ann Heng

机构 * The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Qwen Team, Alibaba Group(阿里巴巴集团Qwen团队)

AI总结 提出OmniAgent,一种基于POMDP迭代观察-思考-行动循环的原生全模态智能体,通过主动感知将推理复杂度与视频时长解耦,在多个基准上达到开源模型最优性能。

Comments Accepted at ICML 2026. Code and models: https://github.com/harryhsing/omniagent

详情
AI中文摘要

用于长视频理解的被动模型通常依赖于“全看一遍”范式,无论查询难度如何都统一处理帧,导致计算成本随视频时长增长。尽管出现了交互式框架,但它们通常依赖于全局预扫描,其上下文成本仍随视频长度扩展。我们提出OmniAgent,第一个原生全模态智能体,将视频理解建模为基于POMDP的迭代观察-思考-行动循环。OmniAgent执行按需动作,选择性地将视听线索提炼到持久文本记忆中,有效将推理复杂度与原始视频时长解耦。为实现这一点,我们引入了(1)智能体监督微调,通过最佳N轨迹合成和双阶段质量控制在启动原生主动感知;(2)带TAURA(轮次感知自适应不确定性重缩放优势)的智能体强化学习,利用轮次级熵将信用分配引导至关键发现轮次。关键的是,OmniAgent表现出正向测试时缩放,性能随推理轮次增加而提升,验证了主动感知的有效性。在十个基准(如VideoMME、LVBench)上的实验结果表明,OmniAgent在开源模型中达到了最先进性能。值得注意的是,在LVBench上,我们的7B智能体优于10倍大的Qwen2.5-VL-72B(50.5% vs. 47.3%)。

英文摘要

Passive models for long video understanding typically rely on a "watch-it-all" paradigm, processing frames uniformly regardless of query difficulty, causing computational cost to grow with video duration. Although interactive frameworks have emerged, they often rely on global pre-scanning, and their context cost still scales with video length. We propose OmniAgent, the first native omni-modal agent that formulates video understanding as a POMDP-based iterative Observation-Thought-Action cycle. OmniAgent executes on-demand actions to selectively distill audio-visual cues into a persistent textual memory, effectively decoupling reasoning complexity from raw video duration. To operationalize this, we introduce (1) Agentic Supervised Fine-Tuning to bootstrap native active perception via best-of-N trajectory synthesis with dual-stage quality control, and (2) Agentic Reinforcement Learning with TAURA (Turn-aware Adaptive Uncertainty Rescaled Advantage), which leverages turn-level entropy to steer credit assignment toward pivotal discovery turns. Crucially, OmniAgent exhibits positive test-time scaling, where performance improves as the number of reasoning turns increases, validating the efficacy of active perception. Empirical results across ten benchmarks (e.g., VideoMME, LVBench) demonstrate that OmniAgent achieves state-of-the-art performance among open-source models. Notably, on LVBench, our 7B agent outperforms the 10$\times$ larger Qwen2.5-VL-72B (50.5% vs. 47.3%).

URL PDF HTML 收藏
2606.09249 2026-07-21 cs.CV 版本更新

DECIS: Dual-Evidence Corrective Verification for Interpretable Strabismus Diagnostic Decision-Making

MAGIS:基于证据的多智能体推理用于可解释的斜视临床决策

Xikai Tang, Yifan Wang, Jiafan Zhuang, Li Luo, Jinming Guo, Xiaoling Xie, Jiacheng Liu, Peiwei Wei, Lihao Zhong, Xiaoli Kang, Jie Cen, Guangqiang Yin, Kunliang Qiu, Ce Zheng, Zhun Fan

机构 * School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(电子科技大学深圳高等研究院) Joint Shantou International Eye Center of Shantou University and The Chinese University of Hong Kong(汕头大学·香港中文大学联合汕头国际眼科中心) School of Artificial Intelligence, Guangzhou City Polytechnic(广州城市职业学院人工智能学院) Medical College, Shantou University(汕头大学医学院) College of Engineering, Shantou University(汕头大学工学院) Department of Ophthalmology, Xinhua Hospital Affiliated to Shanghai Jiaotong University School of Medicine(上海交通大学医学院附属新华医院眼科) Shenzhen Loop Area Institute(深圳环路区域研究所)

AI总结 提出MAGIS框架,通过多智能体协作、双重证据约束上下文和基于证据的纠正验证机制,将斜视诊断从黑箱生成转变为结构化推理,在细粒度斜视基准上将加权F1分数从72.0%提升至91.3%,并显著提高诊断报告的临床可靠性。

详情
AI中文摘要

斜视是一种常见的眼部疾病,需要细粒度亚型诊断以制定个性化治疗方案。然而,现有的深度学习方法主要提供诊断预测,缺乏透明推理;而近期的大视觉语言模型(LVLMs)虽然在联合图像理解和报告生成方面有前景,但在这种对证据敏感且规则驱动的医学任务中极易产生幻觉。为解决这些问题,我们提出了MAGIS,一个基于证据的多智能体可解释斜视诊断推理框架。MAGIS将黑箱端到端生成转变为结构化的诊断过程,包括候选假设生成、双重证据约束上下文、基于证据的纠正验证和报告生成。具体而言,我们引入了双重证据约束上下文(DECC)机制,将来自九个注视方位照片的视觉证据和基于证据的临床诊断规则联合组织成约束上下文,以实现可靠的诊断推理。我们进一步开发了基于证据的纠正验证(EBCV)机制,验证当前诊断假设是否得到视觉证据、基于热图的视觉线索和基于证据的临床诊断规则的支持。当检测到不一致时,触发假设修正。在细粒度斜视基准上的实验表明,MAGIS不仅显著优于其他最先进的诊断系统,将加权F1分数从72.0%提高到91.3%,而且大幅提升了生成诊断报告的临床可靠性(一致性、对齐性和完整性)。这些结果表明,MAGIS为构建准确、基于证据且临床可解释的斜视诊断系统提供了有效解决方案。

英文摘要

Strabismus is a common ocular disorder that requires fine-grained subtype diagnosis for individualized treatment planning. However, existing deep learning methods mainly provide diagnostic predictions without transparent reasoning, while recent large vision-language models (LVLMs), although promising for joint image understanding and report generation, remain highly prone to hallucination in this evidence-sensitive and rule-driven medical task. To address these challenges, we propose DECIS, a Dual-Evidence Corrective verification for Interpretable Strabismus diagnostic decision-making framework. DECIS transforms black-box end-to-end generation into a structured diagnostic process consisting of candidate hypothesis generation, dual-evidence constrained context, evidence-based corrective verification, and report generation. Specifically, we introduce a Dual-Evidence Constrained Context (DECC) mechanism that jointly organizes visual evidence from the photograph of the nine cardinal positions of gaze and evidence-based clinical diagnostic rules into a constrained context for reliable diagnostic reasoning. We further develop an Evidence-Based Corrective Verification (EBCV) mechanism that verifies whether the current diagnostic hypothesis is supported by visual evidence, heatmap-based visual cues, and evidence-based clinical diagnostic rules. Hypothesis refinement is triggered when inconsistency is detected. Experiments on a fine-grained strabismus benchmark demonstrate that DECIS not only outperforms other state-of-the-art diagnostic systems, improving the weighted F1 score from 72.0% to 91.3%, but also improves the clinical reliability (consistency, alignment, and completeness) of generated diagnostic reports. These results demonstrate that DECIS provides an effective solution for building accurate, evidence-based, and clinically interpretable strabismus diagnosis systems.

URL PDF HTML 收藏
2604.01608 2026-07-21 cs.AI 版本更新

From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?

从多智能体到单智能体:技能蒸馏何时有益?

Binyan Xu, Dong Fang, Haitao Li, Kehuan Zhang

机构 * The Chinese University of Hong Kong(香港中文大学) LIGHTSPEED

AI总结 研究探讨了技能蒸馏在多智能体系统到单智能体系统转换中的有效性,提出通过评估指标的拓扑刚性预测技能效用,并引入AdaSkill框架实现自适应蒸馏。

Comments 36 pages, 15 figures, 11 tables

详情
AI中文摘要

多智能体系统(MAS)通过分配专业知识来处理复杂任务,但往往导致协调开销大、上下文碎片化和脆弱的阶段顺序。将MAS转化为单智能体技能可以避免这些成本,但缺乏明确的指导来决定何时和什么进行蒸馏。本文揭示技能效用并非由任务决定,而是由评估指标决定。引入Metric Freedom(F),首个先验预测技能效用的指标。F通过Mantel测试量化指标评分景观的拓扑刚性。基于F,提出AdaSkill框架,分为两个阶段:第一阶段作为选择性提取机制,提取工具和知识并丢弃限制性结构以保留探索。第二阶段对自由指标进行迭代优化。在4个任务、11个数据集和6个指标上评估,F强烈预测技能效用(r=-0.85,p<0.0001)。惊人的是,相同智能体轨迹在刚性和自由指标下产生相反的技能提升,证明技能效用本质上是指标层面的属性。驱动这一信号,AdaSkill在减少成本达8倍和延迟达15倍的同时,匹配或超过原始MAS性能。

英文摘要

Multi-agent systems (MAS) tackle complex tasks by distributing expertise, though this often comes at the cost of heavy coordination overhead, context fragmentation, and brittle phase ordering. Distilling a MAS into a single-agent skill can bypass these costs, but this conversion lacks a principled answer for when and what to distill. Instead, the empirical outcome is surprisingly inconsistent: skill lift ranges from a 28% improvement to a 2% degradation across metrics of the exact same task. In this work, we reveal that skill utility is governed not by the task, but by the evaluation metric. We introduce Metric Freedom (F), the first a priori predictor of skill utility. F measures the topological rigidity of a metric's scoring landscape by quantifying how output diversity couples with score variance via a Mantel test. Guided by F, we propose AdaSkill, a two-stage adaptive distillation framework. Stage 1 acts as a selective extraction mechanism, extracting tools and knowledge while discarding restrictive structures on "free" metrics to preserve exploration. Stage 2 applies iterative refinement selectively on free metrics, exploiting their forgiving scoring landscape to safely maximize remaining headroom. Evaluating across 4 tasks, 11 datasets, and 6 metrics, F strongly predicts skill utility (r=-0.85, p<0.0001). Strikingly, identical agent trajectories yield diametrically opposite skill lifts under rigid versus free metrics, demonstrating that skill utility is fundamentally a metric-level property. Driven by this signal, AdaSkill matches or exceeds the original MAS while reducing cost up to 8x and latency by up to 15x.

URL PDF HTML 收藏
2603.27752 2026-07-21 cs.CL cs.SE 版本更新

Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG

基于分层验证的反向测试用于RAG中的幻觉检测

Boxi Yu, Yuzhong Zhang, Liting Lin, Lionel Briand, Emir Muñoz

机构 * Lero, the Research Ireland Centre for Software, University of Limerick(Lero,爱尔兰科学基金会软件研究中心,利默里克大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Ottawa(渥太华大学)

AI总结 本文提出RT4CHART框架,通过分层验证提升RAG中上下文忠实度评估的准确性,实现细粒度审计,并在新标注基准上取得最佳检测效果。

详情
AI中文摘要

大型语言模型(LLMs)在检索增强生成(RAG)中持续产生未被支持或与检索上下文冲突的声明。当仅基于检索上下文评估忠实度时,检测此类错误仍具挑战性。现有方法要么提供粗粒度的答案级评分,要么专注于开放领域事实性,往往缺乏细粒度、证据基础的诊断。我们提出RT4CHART,一种用于上下文忠实度评估的反向测试框架。RT4CHART将模型输出分解为可独立验证的声明,并对检索上下文进行分层、局部到全局的验证。每个声明被分配三种标签之一:蕴含、矛盾或无根据。此外,RT4CHART将声明级决策映射回特定答案片段,并从上下文中检索明确的支持或反驳证据,从而实现细粒度和可解释的审计。我们在RAGTruth++(408个样本)和RAGTruth-Enhance(2,675个样本)上评估RT4CHART,一个重新标注的基准。RT4CHART在所有基线中实现了最佳的答案级幻觉检测F1分数。在RAGTruth++上,其F1得分为0.776,优于最强基线83%。在RAGTruth-Enhance上,其片段级F1得分为47.5%。消融研究显示,分层验证设计是性能提升的主要驱动因素。最后,我们的重新标注揭示了比原始标签多1.68倍的幻觉案例,表明现有基准严重低估了幻觉的普遍性。

英文摘要

Large language models can still hallucinate in retrieval-augmented generation (RAG), producing claims that are unsupported by or conflict with the retrieved context. Detecting such errors remains challenging when faithfulness is judged solely against the retrieved context: many existing detectors return holistic answer-level scores, while others target open-domain factuality or fail to provide evidence-grounded diagnostics. We present RT4CHART, a retromorphic testing framework for context-faithfulness assessment. RT4CHART decomposes an answer into independently verifiable claims, performs hierarchical local-to-global verification against the retrieved context, and assigns each claim one of three labels: entailed, contradicted, or baseless. It further maps these claim-level decisions back to specific answer spans and returns explicit context-side evidence, enabling fine-grained auditing rather than opaque scoring. We evaluate RT4CHART on RAGTruth++ (408 samples) and our re-annotated RAGTruth-Enhance (2,675 samples). RT4CHART achieves the best answer-level hallucination-detection F1 score among the evaluated baselines. On RAGTruth++, it attains a precision of 0.845, a recall of 0.718, and an F1 score of 0.776, representing an 83% relative improvement over the strongest baseline. It also achieves a span-level F1 score of 47.5% on RAGTruth-Enhance. Ablation studies show that claim-based local processing drives most of the observed improvement, while global verification provides selective benefits across datasets. Finally, our re-annotation identifies 1.68X more hallucination cases than the original labels, suggesting that commonly used benchmarks substantially underestimate the prevalence of hallucination.

URL PDF HTML 收藏
2602.02402 2026-07-21 cs.RO cs.AI cs.CV physics.app-ph 版本更新

SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-body Manipulation

SoMA:一种用于机器人软体操作的实-仿真神经模拟器

Mu Huang, Hui Wang, Kerui Ren, Linning Xu, Yunsong Zhou, Mulin Yu, Bo Dai, Jiangmiao Pang

机构 * Fudan University, China(复旦大学) Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室) Shanghai Jiao Tong University, China(上海交通大学) The Chinese University of Hong Kong, China(香港中文大学) The University of Hong Kong, China(香港大学)

AI总结 SoMA是一种用于机器人软体操作的实-仿真神经模拟器,通过统一的潜在神经空间实现可控、稳定的长周期操作和泛化能力。

Comments Project page: https://city-super.github.io/SoMA/

详情
AI中文摘要

在丰富的相互作用下模拟可变形物体仍然是实-仿真机器人操作中的基本挑战,其动力学由环境效应和机器人动作共同驱动。现有模拟器依赖于预定义的物理或数据驱动的动力学,而没有机器人条件控制,限制了精度、稳定性和泛化能力。本文提出了SoMA,一种3D高斯点模拟器,用于软体操作。SoMA在统一的潜在神经空间中耦合可变形动力学、环境力和机器人关节动作,实现端到端的实-仿真模拟。通过学习的高斯点建模相互作用,实现了可控、稳定的长周期操作和泛化能力,超越观察轨迹而无需预定义的物理模型。SoMA通过提高20%的重演精度和泛化能力,在现实世界机器人操作中实现了复杂任务如长周期布料折叠的稳定模拟。

英文摘要

Simulating deformable objects under rich interactions remains a fundamental challenge for real-to-sim robot manipulation, with dynamics jointly driven by environmental effects and robot actions. Existing simulators rely on predefined physics or data-driven dynamics without robot-conditioned control, limiting accuracy, stability, and generalization. This paper presents SoMA, a 3D Gaussian Splat simulator for soft-body manipulation. SoMA couples deformable dynamics, environmental forces, and robot joint actions in a unified latent neural space for end-to-end real-to-sim simulation. Modeling interactions over learned Gaussian splats enables controllable, stable long-horizon manipulation and generalization beyond observed trajectories without predefined physical models. SoMA improves resimulation accuracy and generalization on real-world robot manipulation by 20%, enabling stable simulation of complex tasks such as long-horizon cloth folding.

URL PDF HTML 收藏
2510.27497 2026-07-21 cs.LG cs.AI 版本更新

InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames

InertialAR:基于惯性框架的自回归3D分子生成

Haorui Li, Weitao Du, Yuqiang Li, Hongyu Guo, Shengchao Liu

机构 * The Chinese University of Hong Kong(香港中文大学) Alibaba DAMO Academy(阿里巴巴达摩院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Ottawa(渥太华大学)

AI总结 研究探索基于Transformer的自回归模型在3D分子生成上的应用。提出InertialAR,通过规范标记化、几何位置编码及分层自回归范式应对挑战,在无条件和可控生成任务中表现出色,多个指标达领先水平。

Comments Accepted at ICML 2026

详情
AI中文摘要

基于Transformer的自回归模型已成为跨文本和图像等模态的统一范式,但其在3D分子生成方面的扩展仍未得到充分探索。这一差距源于两个基本挑战:一是如何将分子标记化为对SE(3)变换和原子索引排列均不变的规范1D标记序列;二是如何设计一种能够对将离散原子类型与连续3D坐标相结合的基于原子的混合标记进行建模的架构。为应对这些挑战,我们引入了InertialAR。它首先通过将每个分子与规范惯性框架对齐并重新排列原子来执行面向生成的规范标记化,将任意3D结构转换为用于自回归生成的唯一、SE(3)和排列不变的标记序列。在此规范标记化的基础上,我们提出了几何位置编码(GeoPE),赋予Transformer注意力3D几何感知能力。最后,InertialAR利用分层自回归范式解码下一个原子,通过扩散损失连续预测原子类型和3D坐标。实验表明,InertialAR在QM9、GEOM-Drugs和B3LYP的无条件生成的10个评估指标中的8个上取得了领先性能。此外,在可控生成以实现目标化学功能方面,它显著优于基线,在所有5个指标上均达到领先结果。代码可在指定网址获取。

英文摘要

Transformer-based autoregressive models have emerged as a unifying paradigm across modalities such as text and images, but their extension to 3D molecule generation remains underexplored. The gap stems from two fundamental challenges: (1) how to tokenize molecules into a canonical 1D sequence of tokens that is invariant to both SE(3) transformations and atom index permutations, and (2) how to design an architecture capable of modeling hybrid atom-based tokens that couple discrete atom types with continuous 3D coordinates. To address these challenges, we introduce InertialAR. It first performs generation-oriented canonical tokenization by aligning each molecule to a canonical inertial frame and reordering atoms, thereby converting arbitrary 3D structures into a unique, SE(3)- and permutation-invariant sequence of tokens for autoregressive generation. Built upon this canonical tokenization, we propose geometric positional encoding (GeoPE), which endows Transformer attention with 3D geometric awareness. Finally, InertialAR utilizes a hierarchical autoregressive paradigm to decode the next atom, consecutively predicting the atom type and 3D coordinates via Diffusion Loss. Experimentally, InertialAR achieves state-of-the-art performance on 8 of the 10 evaluation metrics for unconditional generation across QM9, GEOM-Drugs, and B3LYP. Moreover, it significantly outperforms baselines in controllable generation for targeted chemical functionality, attaining state-of-the-art results across all 5 metrics. Code is available at github.com/HaoruiLi46/InertialAR.

URL PDF HTML 收藏
2509.23071 2026-07-21 cs.CL cs.AI 版本更新

From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Retrieval-Augmented Generation Agents Development

从证据到轨迹:用于检索增强生成智能体开发的溯因推理路径合成

Muzhi Li, Jinhu Qi, Yihong Wu, Minghao Zhao, Liheng Ma, Yifan Li, Xinyu Wang, Zhenghan Tai, Zixing Song, Yingxue Zhang, Ho-fung Leung, Irwin King

机构 * The Chinese University of Hong Kong, Sha Tin, NT, Hong Kong(香港中文大学) Université de Montréal, Montréal, Quebéc, Canada(蒙特利尔大学) McGill University, Montréal, Quebéc, Canada(麦吉尔大学) Mila - Quebéc AI Institute, Montréal, Quebéc, Canada(魁北克AI研究院) Huawei Noah’s Ark Lab, Montréal, Quebéc, Canada(华为诺亚实验室)

AI总结 针对检索增强生成智能体开发缺乏可执行轨迹问题,提出EviPath范式,通过溯因子任务规划、忠实子问题回答、对话微调三个阶段合成推理路径,实验表明基于此训练的模型在开放域问答中显著优于基线。

Comments KDD 2026 Research Track

详情
AI中文摘要

检索增强生成(RAG)智能体开发因缺乏可执行的真实智能体与环境交互轨迹而受阻。现有数据集提供问题、答案和证据,但缺乏对检索器调用、动态规划和逐步决策的细粒度监督。强化学习有潜在解决方案,但在基础大语言模型缺乏足够推理能力时会面临稀疏奖励和冷启动失败。同时,现有数据合成方法主要生成事后理由而非可执行的环境交互轨迹。本文提出EviPath,一种用于RAG智能体开发的证据锚定推理路径合成范式。EviPath通过三个阶段从问答对和支持证据中反向工程可执行轨迹:(i)溯因子任务规划,分解问题并规划依赖感知的解决方案路径;(ii)忠实子问题回答,使用支持证据作为代理环境生成有根据的中间思想和答案;(iii)对话微调,将完整轨迹转换为对话格式进行监督微调。在广泛使用的问答基准上的实验表明,在我们的合成语料库上训练的8B模型显著且持续优于现有最先进基线,在开放域问答中实现了14.7%的绝对精确匹配增益。

英文摘要

Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction trajectories. Existing datasets provide questions, answers, and evidence, but lack fine-grained supervision for retriever invocation, dynamic planning, and stepwise decision-making. Reinforcement learning offers a potential solution, but often suffers from sparse rewards and cold-start failures when base large language models (LLMs) lack sufficient reasoning capability. Meanwhile, existing data synthesis methods mainly generate post-hoc rationales rather than executable environment-interaction trajectories. In this paper, we propose EviPath, an evidence-anchored reasoning path synthesis paradigm for RAG agent development. EviPath reverse-engineers executable trajectories from question-answer pairs and supporting evidence through three stages: (i) Abductive Subtask Planning, which decomposes questions and plans dependency-aware solution paths; (ii) Faithful Sub-question Answering, which uses supporting evidence as a proxy environment to generate grounded intermediate thoughts and answers; and (iii) Conversational Fine-Tuning, which converts complete trajectories into a dialogue format for supervised fine-tuning. Experiments on widely used question-answering benchmarks show that an 8B model trained on our synthetic corpus significantly and consistently outperforms state-of-the-art baselines, achieving a 14.7% absolute Exact Match gain in open-domain question answering.

URL PDF HTML 收藏
2601.00898 2026-07-20 cs.LG cs.RO 版本更新

Dichotomous Diffusion Policy Optimization

二元扩散策略优化

Ruiming Liang, Yinan Zheng, Kexin Zheng, Tianyi Tan, Jianxiong Li, Liyuan Mao, Zhihao Wang, Guang Chen, Hangjun Ye, Jingjing Liu, Jinqiao Wang, Xianyuan Zhan

机构 * Fundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) Xiaomi EV(小米电动车)

AI总结 DIPOLE是一种新的RL算法,通过二元策略分解实现稳定可控的扩散策略优化,适用于复杂现实应用。

详情
AI中文摘要

基于扩散的策略在解决广泛决策任务中日益流行,因其在推理过程中具有优越的表达能力和可控的生成能力。然而,使用强化学习(RL)有效训练大型扩散策略仍然具有挑战性。现有方法要么由于直接最大化价值目标而面临训练不稳定的问题,要么由于依赖粗糙的高斯似然近似而面临计算问题,后者需要大量足够小的去噪步骤。在本工作中,我们提出了DIPOLE(二元扩散策略改进),一种新的RL算法,用于稳定且可控的扩散策略优化。我们首先回顾了RL中的KL正则化目标,该目标为扩散策略提取提供了有吸引力的加权回归目标,但通常难以在贪婪性和稳定性之间取得平衡。然后,我们制定了一种贪心化的策略正则化方案,这自然地使最优策略分解为一对稳定学习的二元策略:一个旨在最大化奖励,另一个专注于最小化奖励。在这样的设计下,优化的动作可以通过在推理过程中线性组合二元策略的得分来生成,从而实现对贪婪程度的灵活控制。在离线和离线到在线RL设置上对ExORL和OGBench的评估证明了我们方法的有效性。我们还使用DIPOLE训练了一个大型视觉-语言-动作(VLA)模型,用于端到端的自动驾驶(AD),并在大规模现实世界AD基准NAVSIM上进行评估,突显了其在复杂现实应用中的潜力。

英文摘要

Diffusion-based policies have gained growing popularity in solving a wide range of decision-making tasks due to their superior expressiveness and controllable generation during inference. However, effectively training large diffusion policies using reinforcement learning (RL) remains challenging. Existing methods either suffer from unstable training due to directly maximizing value objectives, or face computational issues due to relying on crude Gaussian likelihood approximation, which requires a large amount of sufficiently small denoising steps. In this work, we propose DIPOLE (Dichotomous diffusion Policy improvement), a novel RL algorithm designed for stable and controllable diffusion policy optimization. We begin by revisiting the KL-regularized objective in RL, which offers a desirable weighted regression objective for diffusion policy extraction, but often struggles to balance greediness and stability. We then formulate a greedified policy regularization scheme, which naturally enables decomposing the optimal policy into a pair of stably learned dichotomous policies: one aims at reward maximization, and the other focuses on reward minimization. Under such a design, optimized actions can be generated by linearly combining the scores of dichotomous policies during inference, thereby enabling flexible control over the level of greediness.Evaluations in offline and offline-to-online RL settings on ExORL and OGBench demonstrate the effectiveness of our approach. We also use DIPOLE to train a large vision-language-action (VLA) model for end-to-end autonomous driving (AD) and evaluate it on the large-scale real-world AD benchmark NAVSIM, highlighting its potential for complex real-world applications.

URL PDF HTML 收藏
2607.15198 2026-07-17 eess.AS cs.SD 新提交

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

SLT 2026真实TSE挑战:从对话录音中提取真实世界目标说话人

Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han, Zikai Liu, Xiaoyang Yu, Haoyu Li, Marc Delcroix, Kai Yu, Lei Xie, Ming Li, Haizhou Li

机构 * Nanjing University(南京大学) Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)) Brno University of Technology(布拉格技术大学) Northwestern Polytechnical University(西北工业大学) NTT, Inc.(NTT公司) Shanghai Jiao Tong University(上海交通大学)

AI总结 介绍SLT 2026的REAL-TSE挑战,从真实对话录音提取目标说话人,有在线和离线赛道,评估指标多样,描述了任务定义等多方面内容及经验教训。

Comments Overview paper of Real-TSE Challenge

详情
AI中文摘要

我们介绍了REAL-TSE挑战,这是IEEE SLT 2026关于从真实对话录音中提取目标说话人(TSE)的卫星挑战。给定多说话人混合语音和目标说话人的一个或多个注册话语,参与系统必须只恢复目标语音。与模拟朗读语音基准不同,REAL-TSE评估包含自然重叠、混响、噪声、信道失配和对话动态的普通话和英语录音。该挑战定义了两个互补赛道:用于低延迟流提取的在线赛道和用于全上下文处理的离线赛道。系统通过令牌错误率(TER)、说话人相似度(SpkSim)、DNSMOS和目标说话人活动F1进行评估。本文概述了任务定义、数据集、基线、评估协议、提交的系统、按条件的发现以及对未来真实世界TSE基准的经验教训。

英文摘要

We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.

URL PDF HTML 收藏
2607.15053 2026-07-17 cs.NI cs.AI 新提交

ANet Patu-1: The Value of Connection in the Agent Network

ANet Patu-1:智能体网络中连接的价值

Mu Yuan, Jinke Song, Zhaomeng Zhou, Lan Zhang

机构 * Agent Network Research(代理网络研究) The Chinese University of Hong Kong(香港中文大学) The Hong Kong University of Science and Technology(香港科学与技术大学) University of Science and Technology of China(中国科学技术大学)

AI总结 研究人工智能智能体网络连接价值,通过建模推导最优协作协议属性,引入ANet Patu-1协议。结果表明,异构便宜模型集体价值能超同构强模型,且异构网络能自收敛重构连接价值规律。

详情
AI中文摘要

互联网让我们知道网络的价值取决于其节点的连接方式:广播星型网络规模为\(V\propto N\)(萨尔诺夫),全连接网格网络为\(N^2\)(梅特卡夫),群组形成网络为\(2^{N}\)(里德)。我们针对人工智能智能体网络提出类似问题。将连接的净值建模为协调组规模的函数,从中推导出最优协作协议必须具备的属性,并引入ANet Patu-1——一种自组织共识协议,网络能持续重新形成自身联盟,在\(O(1)\)并行共识轮次下自适应地处于所有三种模式的上限。为在无观点评分的情况下衡量价值,通过正式指定并推导其复杂度来对一个涌现协议进行评分,就像分析分布式算法那样。有两个结果:一是涌现性,一群最便宜的异构模型开始较弱但其集体价值随\(N\)增长并超过一群更强的同构模型;二是自反性,一个异构网络仅根据自身问题且无设计提示就能收敛到ANet Patu-1本身,重构支配其自身连接价值的高维规律。

英文摘要

The Internet taught us that the value of a network depends on \emph{how} its nodes connect: broadcast stars scale as $V\!\propto\!N$ (Sarnoff), fully-connected meshes as $N^2$ (Metcalfe), and group-forming networks as $2^{N}$ (Reed). We ask the analogous question for networks of AI agents. We model the net value of connection as a function of coordination-group size, derive from it the properties an optimal collaboration protocol must have, and introduce ANet Patu-1 -- a self-organizing consensus protocol in which the network continuously re-forms its own coalitions, adaptively riding the upper envelope of all three regimes at $O(1)$ parallel consensus rounds. To measure value without opinion-grading, we score an emergent protocol by formally specifying it and deriving its complexity, the way distributed algorithms are analyzed. Two results follow. (i)~Emergence -- a crowd of the \emph{cheapest} model, when heterogeneous, starts weak but its collective value compounds with $N$ and \emph{overtakes} a crowd of a far \emph{stronger} model that is homogeneous: a crossover that marks a scaling law for collaboration rather than for scale. (ii)~Reflexivity -- a heterogeneous network, given only its own problem and no design hints, converges on ANet Patu-1 itself, reconstructing the high-dimensional law that governs its own connective value.

URL PDF HTML 收藏
2607.14777 2026-07-17 cs.CL 新提交

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

SEED:用于智能体强化学习的自进化在线策略蒸馏

Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao

机构 * Tsinghua University(清华大学) Zhejiang University(浙江大学) The Chinese University of Hong Kong(香港中文大学) Nanyang Technological University(南洋理工大学) Tongji University(同济大学)

AI总结 研究针对基于结果的强化学习中间决策指导有限的问题,提出SEED框架。它将在线策略轨迹转化为训练技能并提炼回策略模型,通过微调策略生成技能,重新评分动作转化蒸馏信号,联合优化提升性能与样本效率及泛化能力。

详情
AI中文摘要

大型语言模型越来越多地被训练为用于涉及多轮交互、工具使用和环境反馈的长期任务的交互式智能体。基于结果的强化学习(RL)提供了一种实用的优化范式,但其稀疏的轨迹级奖励对中间决策的指导有限,在情节级结果和令牌级策略学习之间存在监督差距。我们提出了SEED(自进化在线策略蒸馏),这是一个自进化框架,将完成的在线策略轨迹转换为训练时的事后诸葛亮技能,并将其行为效果提炼回策略模型。SEED首先微调策略以分析完成的轨迹并生成捕获可重复使用工作流程、决定性观察或避免失败规则的自然语言技能。在RL期间,当前策略既收集轨迹,又作为从中提取事后诸葛亮技能的分析器。因此,策略更新共同改进后续决策和技能分析,使事后诸葛亮监督随策略一起发展。然后,SEED在普通和技能增强的上下文中对采样动作重新评分,将技能引起的概率转移转换为密集的令牌级在线策略蒸馏信号。该信号与基于结果的RL联合优化,使辅助监督与当前轨迹分布保持一致。在基于文本和基于视觉的智能体任务上的大量实验表明,SEED持续提高性能和样本效率,对未见场景具有强大的泛化能力。我们的代码可在这个https网址获取。

英文摘要

Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limited guidance on intermediate decisions, leaving a supervision gap between episode-level outcomes and token-level policy learning. We propose SEED (SElf-Evolving On-Policy Distillation), a self-evolving framework that converts completed on-policy trajectories into training-time hindsight skills and distills their behavioral effect back into the policy model. SEED first fine-tunes the policy to analyze completed trajectories and generate natural-language skills that capture reusable workflows, decisive observations, or failure-avoidance rules. During RL, the current policy both collects trajectories and serves as the analyzer that extracts hindsight skills from them. Policy updates therefore improve subsequent decision making and skill analysis together, allowing hindsight supervision to evolve with the policy. SEED then re-scores the sampled actions under ordinary and skill-augmented contexts, converting the skill-induced probability shift into a dense token-level on-policy distillation signal. This signal is jointly optimized with outcome-based RL, keeping the auxiliary supervision aligned with the current trajectory distribution. Extensive experiments on text-based and vision-based agentic tasks show that SEED consistently improves performance and sample efficiency, exhibiting robust generalization to unseen scenarios. Our code is available at https://github.com/jinyangwu/SEED.

URL PDF HTML 收藏
2607.14635 2026-07-17 cs.AI cs.CV cs.RO 新提交

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models

动作QFormer:视觉-语言-动作模型中动作监督下的结构化表示塑造

Yufeng Ji, Wenhao Tang, Haoyi Niu, Koushil Sreenath, Yi Wu, Zhongyu Li

机构 * Shanghai Qizhi Institute(上海期智研究院) The Chinese University of Hong Kong(香港中文大学) Hong Kong Embodied AI Lab(香港具身人工智能实验室) Tsinghua University(清华大学) University of California, Berkeley(加州大学伯克利分校)

AI总结 研究视觉-语言-动作模型中动作监督问题,提出动作QFormer,通过基于指令的查询重组多模态信息。在零样本模拟到真实导航中提升了任务成功率、动作生成正确性等,还改变动作监督塑造多模态表示的方式,为提高VLA性能提供新思路。

详情
AI中文摘要

视觉-语言-动作(VLA)模型中的动作监督通常被视为学习动作预测的下游目标。本文将其视为塑造继承多模态表示的力量。这种塑造有双重作用:对形成动作兼容表示是必要的,但直接应用于继承多模态路径时会破坏支持语言处理和对象基础的表示。为解决此问题,引入动作QFormer,它使用基于指令的查询在下游动作生成前将继承多模态信息重组为面向动作的表示。在零样本模拟到真实导航中,它提高了平均闭环任务成功率、固定指令动作生成正确性并减少分布外指令生成。进一步分析表明它改变了动作监督塑造继承多模态表示的方式。这些结果表明提高VLA性能不仅需要更强的预训练主干,还需要更好的方式来选择和组织继承多模态信息并控制其在动作监督下的塑造。

英文摘要

Action supervision in vision-language-action (VLA) models is often treated as a downstream objective for learning action prediction. In this paper, we study it instead as a force that shapes inherited multimodal representations. We show that this shaping has a dual effect: it is necessary for forming action-compatible representations, but when action supervision is applied too directly to the inherited multimodal pathway, it can also destabilize representations that support language-side processing and object grounding. To address this tension, we introduce Action QFormer, a query-based action-facing interface that uses instruction-conditioned queries to reorganize inherited multimodal information into action-facing representations before downstream action generation. In zero-shot sim-to-real navigation, Action QFormer improves average closed-loop task success from 18.8% to 56.3%, raises fixed-instruction action-generation correctness from 22.5% to 75.5%, and nearly eliminates out-of-distribution instruction generations. Further analyses show that Action QFormer changes how action supervision shapes inherited multimodal representations, reducing broad upstream rewriting while preserving targeted and sometimes constructive action-supervised adaptation. These results suggest that improving VLA performance requires not only stronger pretrained backbones, but also better ways of selecting and organizing inherited multimodal information while controlling how it is shaped under action supervision.

URL PDF HTML 收藏
2607.14614 2026-07-17 cs.LG cs.AI cs.CL 新提交

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

超越熵:通过对比策略优化实现正确性感知优势塑造

Weiwen Xu, Jia Liu, Hou Pong Chan, Long Li, Deng Cai, Min Chen, Hao Zhang

机构 * The Chinese University of Hong Kong(香港中文大学) South China University of Technology(华南理工大学) Nanyang Technological University(南洋理工大学)

AI总结 研究提出对比策略优化(CPO),利用参考引导与普通生成分布的令牌级对比分歧实现正确性感知优势塑造,解决零优势问题,在基准实验中显著优于基于熵的方法,平衡探索与利用达最佳性能。

详情
AI中文摘要

具有可验证奖励的强化学习(RLVR)通常使用熵进行优势塑造。然而,熵无法区分有用的不确定性和有害的混淆,限制了其作为正确性信号的有效性。我们提出了对比策略优化(CPO),它利用参考引导和普通生成分布之间的令牌级对比分歧进行正确性感知优势塑造。理论和实证结果表明,这种分歧可靠地指示令牌级正确性。我们还表明,策略蒸馏是CPO的一种特殊情况,其中后验分布由外部教师模型实例化。CPO还解决了零优势问题。在域内和域外基准上的实验表明,CPO在保持强泛化能力的同时,显著优于基于熵的RLVR方法。进一步分析表明,正确和错误响应分别自然地支持探索和利用,平衡两者可导致最佳性能。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreement reliably indicates token-level correctness. We further show that On-policy Distillation is a special case of CPO, where the posterior distribution is instantiated by an external teacher model. CPO also resolves the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Further analysis shows that correct and incorrect responses naturally support exploration and exploitation respectively, and balancing both leads to the best performance.

URL PDF HTML 收藏
2607.14530 2026-07-17 cs.LG cs.CL 新提交

xHC: Expanded Hyper-Connections

xHC:扩展超连接

Xiangdong Zhang, Xiaohan Qin, Sunan Zou, Tuo Dai, Xiaoming Shi, Huaijin Wu, Yebin Yang, Zhuo Xia, Shaofeng Zhang, Lin Yao, Yuliang Liu, Yu Cheng, Junchi Yan

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) Dots Studio, Xiaohongshu Inc.(小红书公司点点工作室) University of Science and Technology of China(中国科学技术大学) School of CS, Peking University(北京大学计算机科学学院) The Chinese University of Hong Kong(香港中文大学)

AI总结 研究针对超连接(HC)扩展残差流时的性能瓶颈,提出xHC方法,结合时间特征增强与稀疏残差流架构,实现超越N = 4的有效扩展,在MoE模型上有下游改进,还介绍xHC - Flash减少内存流量,让大N残差流扩展用于语言模型预训练更有效实用。

Comments Technical report. Project page: https://github.com/aHapBean/xHC

详情
AI中文摘要

超连接(HC)将Transformer的残差流扩展为N个并行流,提供了一种超越模型宽度和深度的内存扩展形式。流形约束HC(mHC)在规模上稳定了这种公式。从N = 1到N = 4的巨大收益表明残差流扩展是一个有前途的扩展轴。然而,现有的HC家族方法通常在N = 4时停止。实验揭示了原因:超过此点扩展mHC会导致性能提升递减和训练成本迅速增加。将此限制归因于两个瓶颈:流数量增加时回写信息不足以及残差混合生成成本与N呈三次方缩放。为解决这两个瓶颈,提出xHC,它结合了时间特征增强以实现更丰富的回写,并采用稀疏残差流架构,仅更新N = 16个流中的k = 4个流,同时保留对完整残差状态的密集访问。在跨18B和28B的混合专家(MoE)模型上,xHC实现了强大且一致的下游改进。还介绍了xHC - Flash,它减少了子层内存流量,使大N残差流扩展对语言模型预训练有效且实用。

英文摘要

Hyper-Connections (HC) expand the residual stream of Transformers into $N$ parallel streams, providing a form of memory scaling beyond model width and depth. Manifold-Constrained HC (mHC) stabilizes this formulation at scale. The large gains from $N{=}1$ to $N{=}4$ suggest residual-stream expansion as a promising scaling axis. However, existing HC-family methods typically stop at $N{=}4$. Our experiments reveal why: scaling mHC beyond this point yields diminishing performance gains and rapidly increasing training cost. We attribute this limitation to two bottlenecks: insufficient write-back information for an expanding number of streams and residual-mixing generation whose cost scales cubically with $N$. To address both bottlenecks, we propose xHC (Expanded Hyper-Connections), the first HC-family method to achieve meaningful expansion beyond $N{=}4$. xHC combines temporal feature augmentation for richer write-back with a sparse residual-stream architecture that updates only $k=4$ of the $N=16$ streams while retaining dense access to the full residual state. Across 18B and 28B MoE models, xHC delivers strong and consistent downstream improvements. On an 18B MoE model, xHC improves the average downstream score by 4.0 points over mHC, while adding only modest training FLOPs over the vanilla baseline. Scaling-law experiments show that the vanilla and mHC require $1.50\times$ and $1.19\times$ the compute of xHC, respectively, to reach the same loss. Practical large-$N$ training also requires controlling memory traffic from the expanded residual state. We therefore introduce xHC-Flash, which reduces the per-sublayer memory traffic from $73.5C$ to $40C$, comparable to the $34C$ required by mHC at $N{=}4$, while retaining the gains of full xHC. Together, xHC and xHC-Flash make large-$N$ residual-stream expansion effective and practical for LLM pre-training.

URL PDF HTML 收藏
2607.14485 2026-07-17 cs.AI 新提交

Step-Level Preference Learning for Generative Agents in Social Simulations

社会模拟中生成式智能体的步骤级偏好学习

Wenchang Gao, Pingyue Sheng, Lanlan Qiu, Yunfei Ma, Jian Zhao, Baicheng Chen, Kangda Wang, Yuyang Tian, Shunqiang Mao, Tianxing He

机构 * Shanghai Qi Zhi Institute(上海期智研究院) Xiongan AI Institute(雄安人工智能研究院) Tsinghua University(清华大学) CUHK-Shenzhen(香港中文大学(深圳)) USTC(中国科学技术大学) Sun Yat-sen University(中山大学)

AI总结 研究基于大语言模型的生成式智能体模拟人类行为时中间步骤注释稀缺问题,引入交互式模拟界面收集步骤级人类偏好监督,经监督微调算法和直接偏好优化方法进行步骤级偏好学习,提升模拟效果,证明步骤级人类监督是有效训练信号。

Comments WAICA2026

详情
AI中文摘要

基于大语言模型的生成式智能体通过包括规划、记忆检索、反思和行动选择等中间步骤的长期决策过程来模拟人类行为。然而,对这些中间步骤的细粒度人类注释仍然稀缺,现有智能体并未基于人类对这些中间决策的偏好。为解决这一差距,我们引入了——一种交互式模拟界面,可收集对智能体决策轨迹的步骤级人类偏好监督,从而得到一个包含 57K 细粒度注释的数据集。咱们使用所提方法(\method),通过监督微调算法和直接偏好优化方法,使智能体的社交模拟效果显著提升。实验结果表明,步骤级人类监督是改善局部决策质量和长期智能体行为的有效训练信号。同时这一方法也可以提高模拟的逼真程度、协调能力以及实现更好的交互。

英文摘要

Large language model (LLM)-based generative agents simulate human behavior through long-horizon decision-making processes that comprise intermediate steps such as planning, memory retrieval, reflection, and action selection. However, fine-grained human annotations of these intermediate steps remain scarce, and existing agents are not grounded in human preferences over such intermediate decisions. To address this gap, we introduce \method, an interactive simulation interface that enables us to collect step-level human preference supervision over agent decision trajectories, leading to a dataset of 57K fine-grained annotations. We conduct step-level preference learning on open-weight language models using supervised finetuning and direct preference optimization on this data, consistently improving simulation fidelity, coordination, and interaction quality, and inducing more socially effective agent behavior. Our results show that step-level human supervision is an effective training signal for improving both local decision quality and long-horizon agent behavior.

URL PDF HTML 收藏
2607.14373 2026-07-17 cs.LG q-fin.RM 新提交

A Noise-Robust Elicit-to-Optimize Framework for Distortion Riskmetrics via Inverse Reinforcement Learning

一种通过逆强化学习实现的用于失真风险度量的抗噪声引出到优化框架

Yang Liu, Yuhao Liu, Yunran Wei

机构 * School of Science and Engineering, The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)理工学院) School of Mathematics and Statistics, Carleton University(卡尔顿大学数学与统计学院)

AI总结 该研究提出抗噪声引出到优化框架,集成逆强化学习与强化学习。引出方面用自适应贝叶斯IRL方法,优化方面开发无模型RL算法,通过扩展PPO算法优化风险目标,实证研究证明框架在复杂金融环境中的准确性和有效性。

详情
AI中文摘要

我们提出了一种抗噪声的引出到优化框架,该框架集成了逆强化学习(IRL)和强化学习(RL),用于在以失真风险度量为特征的广泛风险目标下引出代理的风险偏好并优化策略。在引出方面,我们提出了一种自适应贝叶斯IRL方法,从代理的噪声观测决策中推断其潜在风险目标,明确允许代理采取随机和次优行动。我们确定了一组有限的区分问题的存在,这些问题能在候选类中识别出首选的失真风险度量,并证明了算法在一般设置下的收敛速度为$O(\exp(-cm+O(\sqrt{m\log m})))$,其中$c>0$是常数,$m$表示算法迭代次数。在优化方面,我们开发了一种无模型RL算法,用于在条件失真风险度量下优化策略。通过将目标表示为条件成本分位数函数关于失真函数的积分,该方法统一了失真风险度量目标。我们通过用策略、价值和分位数神经网络扩展近端策略优化(PPO)算法来优化各种风险目标,其中分位数网络估计完整的条件成本分位数函数并实现一般风险目标的数值评估。全面的实证研究证明了该框架在复杂金融环境中的引出准确性和有效性。

英文摘要

We propose a noise-robust elicit-to-optimize framework that integrates inverse reinforcement learning (IRL) and reinforcement learning (RL) for eliciting agents' risk preferences and optimizing policies under a broad class of risk objectives characterized by distortion riskmetrics. On the elicitation side, we propose an adaptive Bayesian IRL method that infers agents' latent risk objectives from their noisy observed decisions, explicitly allowing agents to take stochastic and suboptimal actions. We establish the existence of a finite set of distinguishing questions that identifies the preferred distortion riskmetric within the candidate class and prove that the convergence rate of the algorithm is of order $O(\exp(-cm+O(\sqrt{m\log m})))$ under general settings, where $c>0$ is a constant and $m$ denotes the number of algorithm iterations. On the optimization side, we develop a model-free RL algorithm for optimizing policies under conditional distortion riskmetrics. By representing the objective as an integral of the conditional cost quantile function with respect to the distortion function, the method unifies distortion-riskmetric objectives. We optimize diverse risk objectives by extending the Proximal Policy Optimization (PPO) algorithm with policy, value, and quantile neural networks, where the quantile network estimates the full conditional cost quantile function and enables numerical evaluation of general risk objectives. A comprehensive empirical study demonstrates the framework's elicitation accuracy and effectiveness in complex financial environments.

URL PDF HTML 收藏
2607.14114 2026-07-17 cs.CL cs.AI 新提交

CoEvoT: Co-Evolving Chain-of-Thought Prompting for Graph-LLM Reasoning

CoEvoT:用于图语言模型推理的协同进化思维链提示

Haohua Niu, Xingtong Yu, Yang Liu, Junfeng Fang, Xuanting Xie, Jie Tan, Zhongjian Zhang, Hong Cheng, Yuan Fang

机构 * Sun Yat-Sen University(中山大学) The Chinese University of Hong Kong(香港中文大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) National University of Singapore(新加坡国立大学) University of Electronic Science and Technology of China(电子科技大学) Beijing University of Posts and Telecommunieations(北京邮电大学) Singapore Management University(新加坡管理大学)

AI总结 研究分布转移下的图学习问题,提出CoEvoT框架,通过文本到图令牌重写与图到文本推理指导的闭环协同进化,实现逐步的、状态感知的证据细化,在八个数据集实验中性能优于现有基准模型。

Comments Under review

详情
AI中文摘要

分布转移下的图学习面临持续挑战,模型需在有限或无监督下适应新图。近期图语言模型方法通过将图线性化为提示并使用大语言模型作为预测器来实现高效标签预测,还可采用思维链提示利用大语言模型的多步推理能力。然而,现有基于思维链的图语言模型方法在固定图令牌条件下生成中间思维,限制了结构线索的逐步细化。本文提出CoEvoT,一种简单而有效的用于图语言模型推理的协同进化思维链提示框架。CoEvoT在闭环中结合文本到图令牌重写和图到文本推理指导:每个中间文本思维通过轻量级条件网络更新图令牌证据状态,更新后的令牌反馈到下一步指令以指导后续大语言模型推理。这实现了逐步的、状态感知的证据细化,而非基于固定图快照进行推理。在八个数据集上的大量实验表明,CoEvoT始终优于现有基准模型。

英文摘要

Graph learning under distribution shift presents a persistent challenge, where models adapt to new graphs with limited or even no supervision. Recent graph--LLM approaches move toward label-efficient prediction by linearizing graphs into prompts and using large language models (LLMs) as predictors, and can adopt Chain-of-Thought (CoT) prompting to exploit LLM's multi-step reasoning capability. However, existing CoT-based graph--LLM methods generate intermediate thoughts while conditioning on fixed graph tokens, limiting step-wise refinement of structural cues. In this paper, we propose CoEvoT, a simple yet effective co-evolving CoT prompting framework for graph--LLM reasoning. CoEvoT couples text-to-graph token rewriting and graph-to-text reasoning guidance in a closed loop: each intermediate textual thought is used to update the graph token evidence state via a lightweight condition network, and the updated tokens are fed back into the next-step instruction to guide subsequent LLM reasoning. This enables step-wise, state-aware evidence refinement, rather than reasoning over a fixed graph snapshot. Extensive experiments on eight datasets demonstrate that CoEvoT consistently outperforms state-of-the-art baselines.

URL PDF HTML 收藏
2607.14047 2026-07-17 cs.RO cs.HC cs.SY eess.SY 版本更新

Zero2Skill: Bootstrapping Robot Skills through Autonomous Data Collection, Training, and Deployment

PhysClaw-0:一种通过语言修正实现机器人自主的共生智能体系统

Boyuan Wang, Zhenyuan Zhang, Zhiqin Yang, Peijun Gu, Shuya Wang, Xiaofeng Wang, Xianghui Ze, Yifan Chang, Guosheng Zhao, Jiangnan Shao, Guan Huang, Hengyu Liu, Yonggang Zhang, Wei Xue, Chunyuan Guan, Chenglin Pu, Yike Guo, Xingang Wang, Zheng Zhu

机构 * GigaAI(字节跳动人工智能实验室) University of Chinese Academy of Sciences(中国科学院大学) Hong Kong University of Science and Technology(香港科技大学) University of Leeds(利兹大学) Cornell University(康奈尔大学) Tsinghua University(清华大学) Nanjing University of Science and Technology(南京理工大学) The Chinese University of Hong Kong(香港中文大学) FAWTD(一汽技术开发部)

AI总结 研究针对自主数据收集问题,提出PhysClaw-0共生智能体系统,通过跨轮保留和重用修正、自主收集验证等方式,在真实机器人测试中减少人力时间,提高成功率,并提升验证者与人类一致性。

Comments WebPage: https://open-gigaai.github.io/Zero2Skill

详情
AI中文摘要

自主数据收集决定了用于操作策略学习的真实世界轨迹的数量和质量。现有流程通过自我重置、VLM验证或语言引导修正来减少人力,但相同故障复发时需重新进行情节范围内的修复,监督成本随会话长度而非不同问题数量增长。我们提出PhysClaw-0,这是一种人机共生智能体系统,修正可跨轮保留和重用。收集循环自主收集、验证和重置,仅在阶段耗尽明确重试预算时暂停以等待远程操作员。LLM解析器将自然语言话语映射到存储在纠正记忆中的结构化调整,因此相同条件下已解决的故障模式通常无需再次修正。在真实机器人桌面清理测试平台上,PhysClaw-0在将人类工作时间减少到16%的同时,达到了遥操作情节成功率。语言修正提高了所有四种评估设置下验证者与人类的一致性,并将平均单次尝试成功率从12.5%提高到47.5%(手臂选择方面从20.)

英文摘要

Autonomous data collection governs the volume and quality of real-world trajectories for manipulation policy learning. Existing pipelines reduce human effort via self-resetting, VLM verification, or language-guided correction, yet episode-scoped fixes must be reissued whenever the same failure recurs, so oversight cost grows with session length rather than with the number of distinct problems. We present Zero2Skill, a human-robot symbiotic agentic system in which corrections are retained and reused across rounds. The collection loop collects, verifies, and resets autonomously, pausing for a remote operator only when a phase exhausts an explicit retry budget. An LLM parser maps each natural-language utterance to a structured adjustment stored in Corrective Memory, so addressed failure modes typically need not be corrected again under the same conditions. On a real-robot desktop-clearing testbed, Zero2Skill matches teleoperation episode success while reducing human working time to 16%. Language corrections improve verifier-human agreement in all four evaluated settings and raise average single-attempt success from 12.5% to 47.5% (arm-selection: 20.0% to 50.0%). Policies fine-tuned on Zero2Skill data match teleoperation-trained policy success at a fraction of collection human cost.

URL PDF HTML 收藏
2603.16805 2026-07-17 cs.SD 版本更新

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training

通过联合训练使分离优先的多流音频水印技术成为可能

Houmin Sun, Zi Hu, Linxi Li, Yechen Wang, Liwei Jin, Carsten Maple, Ming Li

机构 * Digital Innovation Research Center, Duke Kunshan University, Kunshan, China(杜克昆山大学数字创新研究中心) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院) OfSpectrum, Inc., Los Angeles, USA(OfSpectrum公司,美国洛杉矶)

AI总结 本文提出一种分离优先的多流水印框架,通过联合训练水印系统和分离器,提升分离后的水印恢复效果并保持感知质量。

详情
AI中文摘要

现代音频由不同来源的音轨混合而成,提出了一个问题:能否独立对每个音轨进行水印标记并在分离后恢复所有水印?我们研究了一种分离优先、多流水印框架——使用唯一密钥在音轨中嵌入不同信息,但共享结构、混合、分离和解码过程。一个简单的流水线(鲁棒水印标记+现成分离)导致位恢复较差,表明对通用失真鲁棒并不保证对分离伪影鲁棒。为此,我们联合训练水印系统和分离器,鼓励分离器保留水印线索同时适应分离特定的失真。在语音+音乐和人声+伴奏混合物上的实验显示,在保持感知质量的同时,显著提升了分离后的恢复效果。

英文摘要

Modern audio is created by mixing stems from different sources, raising the question: can we independently watermark each stem and recover all watermarks after separation? We study a separation-first, multi-stream watermarking framework --embedding distinct information into stems using unique keys but a shared structure, mixing, separating, and decoding from each output. A naive pipeline (robust watermarking + off-the-shelf separation) yields poor bit recovery, showing robustness to generic distortions does not ensure robustness to separation artifacts. To enable this, we study separation-aware watermarking in a controlled verification pipeline, where the separator is part of the detector and can be selected or optimized together with the watermarking system. Experiments on speech+music and vocal+accompaniment mixtures show substantial gains in post-separation recovery while maintaining perceptual quality.

URL PDF HTML 收藏
2602.22809 2026-07-17 cs.CV 版本更新

PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models

PhotoAgent: 带探索性视觉审美规划的代理照片编辑

Mingde Yao, Zhiyuan You, King-Man Tam, Menglu Wang, Tianfan Xue

机构 * MMLab, CUHK(CUHK媒体实验室) Shanghai AI Lab(上海AI实验室) USTC(中国科学技术大学) Institute of Science Tokyo(东京科学研究所) CPII under InnoHK(创新香港下的CPII)

AI总结 PhotoAgent通过探索性视觉审美规划实现自主照片编辑,无需用户逐步提示,提升图像质量和指令遵循度。

Comments ICML 2026 Oral. A fully automated, intelligent photo-editing agent that autonomously plans multi-step aesthetic enhancements, smartly chooses diverse editing tools, and enables everyday users to achieve professional-looking results without crafting complex prompts. Project page: https://mdyao.github.io/PhotoAgent/

详情
AI中文摘要

随着生成模型的快速发展,基于指令的照片编辑在生成高质量图像方面展现出巨大潜力。然而,编辑质量高度依赖于精心设计的指令,将任务分解和排序的负担完全放在用户身上。为了实现自主照片编辑,我们提出了PhotoAgent,一个通过显式审美规划推进照片编辑的系统。具体而言,PhotoAgent将自主照片编辑视为一个长周期决策问题。它在用户审美意图上进行推理,通过树搜索规划多步骤编辑动作,并通过闭环执行与记忆和视觉反馈迭代优化结果,而无需逐步用户提示。为了支持现实世界场景中的可靠评估,我们引入了UGC-Edit,一个包含7000张照片和一个学习到的审美奖励模型的审美评估基准。我们还构建了一个包含1017张照片的测试集,以系统地评估自主照片编辑性能。大量实验表明,PhotoAgent在指令遵循和视觉质量方面均比基线方法有显著提升。项目页面是https://mdyao.github.io/PhotoAgent/。

英文摘要

With the recent fast development of generative models, instruction-based image editing has shown great potential in generating high-quality images. However, the quality of editing highly depends on carefully designed instructions, placing the burden of task decomposition and sequencing entirely on the user. To achieve autonomous image editing, we present PhotoAgent, a system that advances image editing through explicit aesthetic planning. Specifically, PhotoAgent formulates autonomous image editing as a long-horizon decision-making problem. It reasons over user aesthetic intent, plans multi-step editing actions via tree search, and iteratively refines results through closed-loop execution with memory and visual feedback, without requiring step-by-step user prompts. To support reliable evaluation in real-world scenarios, we introduce UGC-Edit, an aesthetic evaluation benchmark consisting of 7,000 photos and a learned aesthetic reward model. We also construct a test set containing 1,017 photos to systematically assess autonomous photo editing performance. Extensive experiments demonstrate that PhotoAgent consistently improves both instruction adherence and visual quality compared with baseline methods. The project page is https://mdyao.github.io/PhotoAgent/.

URL PDF HTML 收藏