arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

NVIDIA(英伟达)

至 收录 1122
2607.17454 2026-07-21 cs.RO 新提交

Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation

通过零样本几何评估实现世界行动模型的测试时缩放

Zesen Zhao, Minkyoung Cho, Hui shen, Boyuan Zheng, Kunxiao Gao, Yulong Cao, Z. Morley Mao

机构 * University of Michigan(密歇根大学) NVIDIA(英伟达)

AI总结 研究针对世界行动模型,提出无需训练的选择性测试时缩放框架\methodgated,基于预测未来的跨视图深度重投影一致性排序,经实验验证该方法能提高任务成功率,同时减少额外采样决策点,还识别出相关失败模式。

Comments Extened version of CVPR 2026 EAI workshop

详情
AI中文摘要

测试时缩放通过额外计算改进基础模型推理,但机器人控制需在执行动作前决定额外计算是否有用。世界行动模型(WAMs)使此决策自然化。我们提出了\methodgated,一种用于WAMs的无需训练的选择性测试时缩放框架。首先实例化\method,一个固定预算的最佳N选择器,通过预测未来的跨视图深度重投影一致性对采样的展开进行排序。\methodgated添加了一个轻量级的动作-未来一致性门,仅在初始展开内部不一致时调用\method。在五个基准设置上的实验表明,固定预算的\method在每个设置中都提高了N=8任务的成功率,启用门控后,\methodgated平均恢复了74.8%的始终开启的成功率提升,同时仅在26.2%的决策点触发额外采样。离线诊断表明跨视图重投影是一个强大的无任务标签选择器,我们将错误的低分选择识别为一种失败模式,有助于解释为什么随着N的增加性能会饱和或下降。

英文摘要

Test-time scaling improves foundation-model inference by spending additional computation, but robot control requires deciding whether extra compute is useful before executing an action. World Action Models (WAMs) make this decision natural: each rollout exposes both an action chunk and predicted future observations. We propose \methodgated, a training-free selective test-time scaling framework for WAMs. We first instantiate \method, a fixed-budget Best-of-$N$ selector that ranks sampled rollouts by cross-view depth reprojection consistency of their predicted futures, computed with a frozen geometry foundation model. \methodgated\ adds a lightweight action--future consistency gate that invokes \method\ only when the initial rollout appears internally inconsistent. Across five benchmark--backbone settings on RoboCasa, LIBERO Long, and RoboTwin~2.0, fixed-budget \method\ improves $N{=}8$ task success in every setting, e.g., raising the RoboCasa group average from $66.3\%$ to $68.4\%$ with Cosmos Policy and from $80.8\%$ to $82.5\%$ with X-WAM. With gating enabled, \methodgated\ recovers on average $74.8\%$ of the always-on success gain while triggering additional sampling on only $26.2\%$ of decision points. Offline diagnostics show that cross-view reprojection is a strong task-label-free selector, and we identify false low-score selections as a failure mode that helps explain why performance can saturate or degrade as $N$ increases.

URL PDF HTML 收藏
2607.16610 2026-07-21 cs.AI 新提交

Just A Rather Very Intelligent Spoken Agent

只是一个相当非常智能的语音代理

Chen Chen, Zhehuai Chen

机构 * NVIDIA(英伟达)

AI总结 研究长期人工智能代理与用户交互薄弱问题,引入JarvisBench基准测试,含代理协作和用户交互两轨道,用模块化原型评估,结果显示贾维斯式调解可提升任务性能,其有效性取决于调解器语言模型大脑。

详情
AI中文摘要

长期人工智能代理能力日益增强,但与用户的交互仍很薄弱。多数工作流程中,用户给出初始指令后只能收到部分文本更新,对代理行为及介入时机缺乏清晰认知。当前代理生态系统缺少始终在线的贾维斯式调解器。本文引入JarvisBench基准测试,它包含代理协作和用户交互两个互补轨道,用模块化原型评估,初步结果表明贾维斯式调解可提供基于跟踪的用户问题响应并改善任务性能,有效性很大程度取决于调解器的语言模型大脑。

英文摘要

Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin. In most workflows, users give an initial instruction, receive only selective textual updates, and lose a clear sense of what the agent is doing or when to step in. This leaves a missing part in the current agent ecosystem: an always-on Jarvis-style mediator that keeps the agent continuously reachable to the user. Such a mediator should support real-time spoken interaction with the user, answer questions without interrupting the worker, proactively report progress or confusion, and inject user guidance back into the agent's execution when useful. In this work, we introduce JarvisBench, a benchmark for measuring the dual value of mediation in long-horizon agent workflows. JarvisBench contains two complementary tracks: an agent-collaboration track that measures whether mediation improves downstream task completion, and a user-interaction track that measures whether mediation makes ongoing execution more understandable, responsive, and accessible to users. We instantiate the benchmark with a modular reference Jarvis prototype and evaluate it on 34 text-only WildClaw tasks executed in OpenClaw. Preliminary results with GPT-5.5, Claude Opus 4.7, Gemini-based, and GPT-based worker agents suggest that Jarvis-style mediation can provide trace-grounded responses to user questions and improve task performance when sparse user guidance is injected at appropriate moments. The results also show that effectiveness depends strongly on the mediator's LLM brain, highlighting both the promise of this missing middle layer and the need for broader community effort. Demo page https://cchen1436.github.io/jarvis

URL PDF HTML 收藏
2607.16560 2026-07-21 cs.AI cs.CV cs.LG cs.MM 新提交

From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence

从模态到命题:多模态智能的语言中心框架

Nadine Chang, Maying Shen, Shizhe Diao, Jialiang Wang, Jingde Chen, Thomas Breuel, Pavlo Molchanov, Rafid Mahmood, Jose M. Alvarez

机构 * NVIDIA(英伟达) University of Ottawa(渥太华大学)

AI总结 该研究提出多模态数据语言表示框架,将观察结果表示为原子命题,通过全局语义码本统一为共享词汇表,置于可解释空间,实现跨模态理解等,还在自动驾驶等数据上进行了展示。

详情
AI中文摘要

我们提出了一种用于多模态数据的语言表示,其中任何观察结果,无论是图像、视频还是文本,都被表示为一袋原子命题,即关于场景中实体、动作和关系的简单陈述。全局语义码本将这些统一为规范原子命题的共享词汇表,将每个模态和观察结果置于一个可解释的空间中,该空间跨越细粒度事实到高级概念,并组合成更丰富的概念。这带来了推理的可解释性、跨模态理解和检索以及组合性,从而实现复杂的多模态理解、丰富的数据管理和复杂的结构化检索。我们在自动驾驶和开放世界数据上展示了该框架。

英文摘要

We propose a language representation for multimodal data in which any observation, whether image, video, or text, is expressed as a bag of atomic propositions, simple statements about the entities, actions, and relations in a scene. A global semantic codebook unifies these into a shared vocabulary of canonical atomic propositions, placing every modality and observation into one interpretable space that spans fine grained facts to high level concepts and composes into richer ones. This brings interpretability with reasoning, cross-modal understanding and retrieval, and compositionality that enables complex multimodal understanding, rich data curation and complex structured retrieval. We demonstrate the framework on autonomous driving and open-world data.

URL PDF HTML 收藏
2607.16345 2026-07-21 cs.SE cs.AI cs.LG cs.PF 新提交

AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows

AEVAL:从轶事性到确定性的智能体技能工作流测试

Tejas Singh Anand, Yuet Ying Christina Wang, Wanting Jiang, Steve Masson, Tian Zheng, Bingjie Zhou

机构 * nvidia(NVIDIA公司)

AI总结 研究针对智能体技能工作流测试缺乏确定性和可重复性的问题,提出AEVAL框架,通过执行器与评分器分离等方法,实现确定性、可重复的测试,给出分层修复建议,能将虚假通过率转换为可重复失败信号并记录修复过程。

Comments 8 pages, 1 figure, 1 table

详情
AI中文摘要

现代智能体系统越来越依赖技能,即教大型语言模型智能体执行领域任务的自然语言和代码可安装包。随着技能库增长,开发者需要每次变更的自动化质量信号,但当前评估多是轶事性的,缺乏可重复性和可比性。我们提出AEVAL,一个集成持续集成(CI)的框架,用确定性、可重复的测试管道取代现有做法。每个技能变更触发测试事件,技能在自动化执行器中根据开发者声明的评估契约运行,发出结构化、有证据支持的质量信号供下游CI处理。关键在于执行器和评分器的结构分离,防止智能体在执行中自我纠正并将修补后输出判定为通过这种微妙但普遍的失败模式。我们的贡献包括:(i)具有每个技能契约和每次运行工件模式的确定性、变更触发评估协议;(ii)将自我纠正偏差形式化为朴素智能体评估器的一种独特失败模式;(iii)执行器/评分器分离及首次尝试评分规则和明确的自我纠正跟踪;(iv)作为内联合并请求注释发布的分层、基于证据的修复建议方案(LV1因果,LV2质量)。在多个智能体软件开发工具包的生产智能体堆栈中的实际技能上进行验证,AEVAL将虚假的100%通过率转换为可重复的首次尝试失败信号,并带有每个执行器修复的可审计记录。

英文摘要

Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a domain task. As skill repositories grow, developers need automated quality signals on every change, yet evaluation today is largely anecdotal: a developer asks an agent to "try the skill," watches a demo, and forms a subjective impression. This yields neither reproducibility across runs nor comparability across versions, and scales poorly to marketplaces where one regression can silently break dozens of downstream workflows. We present AEVAL (Agentic Evaluation), a CI-integrated framework that replaces this practice with a deterministic, reproducible test pipeline for agentic skills. Every skill change triggers a test event: the skill runs against a developer-declared evaluation contract (eval.config) inside an automated executor, emitting a structured, evidence-grounded quality signal that downstream CI can route on. A key ingredient is a structural separation between executor and grader, preventing a subtle but pervasive failure mode: an agent that silently self-corrects during execution and then grades its own patched outputs as passing. Our contributions are: (i) a deterministic, change-triggered evaluation protocol with per-skill contracts and per-run artifact schemas; (ii) a formalization of self-correction bias as a distinct failure mode of naive agentic evaluators; (iii) an executor/grader separation with a first-attempt grading rule and explicit self-correction tracking; and (iv) a tiered, grounded-evidence fix-suggestion scheme (LV1 causal, LV2 quality) posted as inline merge-request comments. Validated on real skills in a production agentic stack across multiple agent SDKs, AEVAL converts spurious 100% pass rates into reproducible first-attempt fail signals with an auditable record of every executor fix.

URL PDF HTML 收藏
2607.12463 2026-07-21 cs.AI cs.CL 版本更新

Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

作为编码智能体基础模型中间训练的函数感知中间填充

Yubo Wang, Jiarong Liang, Yuxuan Zhang, Xuye Liu, Cong Wei, Yuyu Zhang, Ping Nie, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学) University of British Columbia(英属哥伦比亚大学) NVIDIA(英伟达公司) Verdent AI(Verdent人工智能公司) Vector Institute(向量研究所)

AI总结 研究针对编码智能体将外部工具返回集成到推理中的问题,利用函数感知中间填充进行中间训练,在多个模型上提升了性能,并减轻了后训练对非智能体编码等基准测试的能力侵蚀。

详情
AI中文摘要

编码智能体必须将外部工具返回结果集成到正在进行的推理中,而标准的代码从左到右预训练仅在正向暴露此能力。我们观察到编码智能体的动作-观察-延续循环在结构上与函数调用站点同构。我们通过函数感知中间填充(FIM)中间训练来利用这一点,这是一种自监督目标,通过程序依赖图分析和复杂性可推断性双重标准来屏蔽函数。我们在从968个GitHub仓库抽取的26亿令牌的净化语料库上对Qwen2.5-Coder-Instruct(7B/14B)和Qwen3-8B进行中间训练,然后应用现有的智能体后训练管道。中间训练在不同模型和后训练管道上都有提升,还减轻了后训练对非智能体编码和非编码工具使用基准测试的能力侵蚀。

英文摘要

Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the action-observation-continuation loop of a coding agent is structurally isomorphic to a function call site, where a caller binds arguments, a callee returns a value computed elsewhere, and downstream code consumes that value. This conditioning structure exists at internet scale in ordinary code. We exploit it through function-aware fill-in-the-middle (FIM) mid-training: a self-supervised objective that masks functions selected via program dependency graph analysis and a complexity-inferability double criterion. We mid-train Qwen2.5-Coder-Instruct (7B/14B) and Qwen3-8B on a 2.6B-token decontaminated corpus drawn from 968 GitHub repositories, then apply existing agentic post-training pipelines. Mid-training improves SWE-Bench-Verified by +2.8/+3.0 at 7B/14B and by +3.2 on Qwen3-8B; SWE-Bench-Lite gains are +3.7/+4.0/+5.4 on the same models. The improvement holds across two post-training pipelines (R2E-Gym, SWE-Smith) and on a non-Qwen2.5 base (Qwen3-8B with SWE-Lego). Beyond in-domain gains, mid-training also mitigates the capability erosion that agentic post-training otherwise inflicts on non-agent coding (e.g., LiveCodeBench) and non-coding tool-use benchmarks (tau-bench, BFCL): although the mid-training corpus contains Python code only, the function-call inductive bias survives post-training and yields consistent gains.

URL PDF HTML 收藏
2607.08601 2026-07-21 cs.CL 版本更新

It Takes a MAESTRO To Prune Bad Experts

需要一位大师来修剪不良专家

Palaash Goel, Ayush Maheshwari, Tanmoy Chakraborty

机构 * Indian Institute of Technology Delhi(印度理工学院德里分校) NVIDIA(英伟达)

AI总结 研究针对稀疏激活的MoE语言模型部署瓶颈问题,提出MAESTRO结构化剪枝框架,将专家激活轨迹建模为马尔可夫链以产生全局感知启发式方法,在多领域评估中性能优于基线,跨任务方差低,模型泛化更一致。

Comments 19 pages, 4 figures

详情
AI中文摘要

稀疏激活的专家混合(MoE)语言模型通过每次只激活一小部分参数来实现显著的推理效率,但其完整的专家库始终驻留在内存中,造成了高昂的部署瓶颈。现有结构化剪枝方法主要针对密集变压器设计,使用局部启发式方法评估专家重要性,而忽视了MoE路由的相互依赖性质。我们引入了MAESTRO(通过基于转移的路由进行马尔可夫链近似专家稀疏化),这是一个为MoE架构设计的结构化剪枝框架,它将自回归专家激活轨迹建模为遍历马尔可夫链,其平稳分布编码跨层依赖性,产生全局感知重要性启发式方法。在包括安全、偏差和伦理在内的五个不同领域进行评估,在严格的50%压缩机制下,MAESTRO在平均性能保留率上比现有基线高出10.61%,同时跨任务方差显著降低,表明全局、路由一致的剪枝产生的模型在异构任务中更一致地泛化。

英文摘要

Sparsely-activated Mixture-of-Experts (MoE) language models achieve remarkable inference efficiency by activating only a small fraction of parameters per token, yet their full expert banks reside in memory at all times, creating a prohibitive deployment bottleneck. Existing structured pruning methods, largely designed for dense transformers, assess expert importance using locally derived heuristics that are blind to the interdependent nature of MoE routing. We introduce MAESTRO (Markov-chain Approximated Expert Sparsification via Transition-based ROuting), a structured pruning framework designed for MoE architectures that models autoregressive expert activation trajectories as Ergodic Markov chains whose stationary distributions encode cross-layer dependencies, yielding a globally aware importance heuristic. Evaluated across five diverse domains including Safety, Bias, and Ethics, MAESTRO outperforms state-of-the-art baselines by up to 10.61% in average performance retention under a strict 50% compression regime, while exhibiting substantially lower cross-task variance, indicating that global, routing-congruent pruning produces models that generalize more consistently across heterogeneous tasks.

URL PDF HTML 收藏
2505.20161 2026-07-21 cs.LG cs.AI cs.CL 版本更新

Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning

棱柱形合成:基于梯度的数据多样化提升语言模型推理中的泛化能力

Jaehun Jung, Seungju Han, Ximing Lu, Skyler Hallinan, David Acuna, Shrimai Prabhumoye, Mostafa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Yejin Choi

机构 * NVIDIA Research(NVIDIA研究部) University of Washington(华盛顿大学) University of Southern California(南加州大学)

AI总结 研究语言模型训练数据多样性对泛化的作用,提出基于梯度熵的G - Vendi指标,进而构建棱柱形合成框架生成多样合成数据,有效提升模型性能,在多个基准测试中表现优于依赖更大数据生成器的模型。

详情
AI中文摘要

语言模型中的有效泛化关键取决于训练数据的多样性。现有多样性指标常依赖与模型行为脱节的表面启发式方法,难以达成目标。为此研究何种训练数据多样性驱动语言模型泛化及如何衡量与增强它。通过超300次训练运行的大规模实证分析表明,数据多样性可有力预测语言模型推理中的泛化。引入G - Vendi指标,基于模型诱导梯度的熵量化多样性,表现优于其他方法。在此基础上提出棱柱形合成框架,通过针对梯度空间中代表性不足区域生成多样合成数据。实验结果显示,随着合成数据规模增加,棱柱形合成持续提升模型性能,在多个基准测试中显著优于依赖更大数据生成器的现有模型。

英文摘要

Effective generalization in language models depends critically on the diversity of their training data. Yet existing diversity metrics often fall short of this goal, relying on surface-level heuristics that are decoupled from model behavior. This motivates us to ask: What kind of diversity in training data actually drives generalization in language models -- and how can we measure and amplify it? Through large-scale empirical analyses spanning over 300 training runs, carefully controlled for data scale and quality, we show that data diversity can be a strong predictor of generalization in LLM reasoning -- as measured by average model performance on unseen out-of-distribution benchmarks. We introduce G-Vendi, a metric that quantifies diversity via the entropy of model-induced gradients. Despite using a small off-the-shelf proxy model for gradients, G-Vendi consistently outperforms alternative measures, achieving strong correlation (Spearman's $ρ\approx 0.9$) with out-of-distribution (OOD) performance on both natural language inference (NLI) and math reasoning tasks. Building on this insight, we present Prismatic Synthesis, a framework for generating diverse synthetic data by targeting underrepresented regions in gradient space. Experimental results show that Prismatic Synthesis consistently improves model performance as we scale synthetic data -- not just on in-distribution test but across unseen, out-of-distribution benchmarks -- significantly outperforming state-of-the-art models that rely on 20 times larger data generator than ours. For example, PrismMath-7B, our model distilled from a 32B LLM, outperforms R1-Distill-Qwen-7B -- the same base model trained on proprietary data generated by 671B R1 -- on 6 out of 7 challenging benchmarks.

URL PDF HTML 收藏
2607.16107 2026-07-20 eess.AS cs.CV 新提交

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

视听火烈鸟:用于长而复杂视频的开放视听智能

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping

机构 * NVIDIA, USA(美国NVIDIA公司) University of Maryland, USA(美国马里兰大学)

AI总结 研究提出视听火烈鸟,用于长复杂视频的联合理解与推理。贡献包括构建大规模视频集、设计三阶段课程及推理框架。实验显示其在多基准测试中表现优异,超越同类开放模型,在长复杂视听理解推理任务中竞争力强,有现实应用价值和泛化能力。

Comments Project Page: https://avflamingo.pages.dev/

详情
AI中文摘要

我们提出了视听火烈鸟(AV - Flamingo),这是一种完全开放的最先进的视听大语言模型(AV - LLM),用于对音频、图像和长视频进行联合理解与推理。与之前主要关注短视频片段的AV - LLM不同,AV - Flamingo旨在对长而复杂的现实世界(视听)视频进行理解和推理。为此,我们做出了三项关键贡献:一是视听技能,一个大规模的现实世界视频集合;二是新颖的三阶段课程;三是时间视听交错思维链推理框架。实验表明AV - Flamingo在多个基准测试中表现出色,具有很强的现实应用价值和泛化能力。

英文摘要

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.

URL PDF HTML 收藏
2607.16094 2026-07-20 cs.CV 新提交

How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA

视觉语言模型如何失败?组合式视觉问答中的视觉-操作不对齐

Navya Gupta, Bingjie Xu, Avinash Anand, Timothy Liu, Zhengchen Zhang

机构 * Singapore Institute of Technology(新加坡科技学院) NVIDIA(英伟达)

AI总结 研究组合式视觉问答中视觉语言模型失败的机制,引入以操作为中心的框架分解失败模式,揭示四种失败模式及传播路径,表明不同失败类型需不同纠正策略,为提升模型可靠性提供基础。

Comments Accepted at ACM Multimedia 2026

详情
AI中文摘要

组合式视觉问答要求视觉语言模型执行多种推理操作,如对象选择、空间关系解析和属性验证。尽管总体性能强劲,但视觉语言模型在此任务上失败的机制基础仍未得到充分探索。为填补这一空白,我们通过研究失败如何与特定推理操作以及它们出现和传播的内部计算路径相关,来分析视觉语言模型中的视觉-操作不对齐。我们引入了一个以操作为中心的机制框架,该框架根据失败产生的推理操作和传播的内部计算路径对视觉语言模型的失败进行分解。我们的分析揭示了四种机制上不同的失败模式:基础失败、推理失败、属性提取失败和语言先验主导失败。每种模式都以视觉基础强度和答案正确性之间的独特关系为特征。通过在所有变压器层应用的三种互补因果干预,我们进一步证明了一种路径分离:基础失败仅通过前馈网络传播,推理失败通过后期层注意力传播,属性提取失败定位到答案位置的前馈计算。这种分离表明不同的失败类型需要根本不同的纠正策略,为有针对性地提高视觉语言模型在多媒体推理中的可靠性提供了原则基础。

英文摘要

Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations like object selection, spatial relation resolution, and attribute verification. Despite strong aggregate performance, the mechanistic basis of VLM failures on this task remains underexplored. To address this gap, we analyze vision-operation misalignment in VLMs by examining how failures relate to specific reasoning operations and the internal computational pathways through which they arise and propagate. We introduce an Operation-centric mechanistic framework that decomposes VLM failures by both the reasoning operation where they originate and the internal computational pathway through which they propagate. Our analysis reveals four mechanistically distinct failure modes: grounding failure, reasoning failure, attribute extraction failure, and language prior dominance failure. Each characterized by a unique relationship between visual grounding strength and answer correctness. Through three complementary causal interventions applied across all transformer layers, we further demonstrate a pathway dissociation: grounding failures route exclusively through the feedforward network, reasoning failures route through late-layer attention, and attribute extraction failures localize to the answer-position feedforward computation. This dissociation demonstrates that different failure types require fundamentally different corrective strategies, providing a principled foundation for targeted improvements to VLM reliability in multimedia reasoning.

URL PDF HTML 收藏
2607.15579 2026-07-20 cs.RO cs.HC 新提交

PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction

PACE:通过人机交互中的对话启发实现角色适应

Peizhen Li, Longbing Cao, Megani Rajendran, Timothy Liu, Aik Beng Ng, Simon See

机构 * Macquarie University(麦考瑞大学) NVIDIA(英伟达)

AI总结 研究如何让类人机器人有可适应角色,提出PACE框架,通过用户问答动态生成角色,经提高结构化提示,经系统集成转化为行为,实证评估显示其对人机交互多方面有积极影响,为部署个性化等身份开辟途径。

Comments 8 pages, 5 figures

详情
AI中文摘要

为类人机器人配备连贯且可适应的角色对于促进自然、引人入胜且值得信赖的人机交互至关重要。然而,现有方法往往依赖缺乏灵活性的静态、硬编码身份。本文提出PACE,一种在Ameca类人机器人上交互式生成和部署结构化角色的新框架。系统引入交互式角色启发管道,通过用户问答动态合成定制且基于心理的身份。该启发过程进入角色提示编译阶段,生成基于多视角维度的结构化角色提示。详细说明了将此结构化规范转化为富有表现力的多模态类人行为所需的实体系统集成。通过全面的人机交互实证评估,与通用基线相比,评估了动态生成角色对用户信任、感知拟人化、角色一致性、个人相关性和交互质量的影响。这些贡献为在实体类人助手部署个性化、交互式和可靠身份建立了可扩展途径。

英文摘要

Equipping humanoid robots with coherent and adaptable personas is crucial for fostering natural, engaging, and trustworthy human-robot interaction (HRI). However, existing approaches often rely on static, hard-coded identities that lack the flexibility to adapt to individual user contexts. In this paper, we present PACE (Persona Adaptation through Conversational Elicitation), a novel framework for the interactive generation and deployment of structured personas on the Ameca humanoid robot. Our system introduces an Interactive Persona Elicitation Pipeline, enabling the robot to dynamically synthesize a tailored, psychologically grounded identity through user Q&A. This elicitation process feeds into a persona prompt compilation phase, generating a structured persona prompt built upon multi-perspective dimensions. We detail the Embodied System Integration required to translate this structured specification into expressive, multimodal humanoid behaviors. Through a comprehensive empirical HRI evaluation, we assess the impact of dynamically generated personas on user trust, perceived anthropomorphism, persona consistency, personal relevance, and interaction quality compared to a generic baseline. These contributions establish a scalable pathway for deploying personalized, interactive, and reliable identities in embodied humanoid assistants. Video demo is available at: https://lipzh5.github.io/PACE/

URL PDF HTML 收藏
2607.15293 2026-07-20 cs.LG cs.AI cs.NA math.NA 新提交

Structure of the Circular-Dyadic Convolution Error

循环二元卷积误差的结构

Ben Fauber, Alireza Moradzadeh

机构 * NVIDIA(英伟达)

AI总结 研究循环二元卷积中用哈达玛变换替代DFT产生的误差,通过确定精确误差抵消、分析误差算子秩及零空间维度、得出期望误差表达式等,揭示该误差有结构、可预测且受对齐控制,除特定子空间滤波器外会使输出能量加倍。

详情
AI中文摘要

二元卷积和循环卷积都可以分别使用哈达玛变换和快速傅里叶变换(FFT)计算的离散傅里叶变换(DFT)在\(O(N\log N)\)时间内完成。哈达玛变换因其实值符号翻转而更受青睐,但其替代DFT会引入代数误差。我们给出了三个互补的结果来刻画这种误差。首先,我们确定了精确的误差抵消:两个输入和两个输出位置普遍无误差,且输出的任何重新排序都无法消除此误差。其次,误差算子几乎是满秩的,而其零空间只有对数维。第三,期望误差由单个对齐标量控制,通过对随机滤波器求平均得到封闭形式的表达式。一般来说,除了通用零误差子空间中的滤波器不会产生误差外,替代误差会使输出能量渐近加倍。总的来说,这些结果表明替代误差是有结构的、可预测的且受对齐控制。

英文摘要

Dyadic and circular convolution can both be computed in $O(N\log N)$ time using the Hadamard transform and the FFT-computed discrete Fourier transform (DFT), respectively. The Hadamard transform is preferable for its real-valued sign flips, yet its substitution for the DFT introduces algebraic error. We present three complementary results that characterize this error. First, we identify exact error cancellation: two input and two output positions are universally error-free, and no reordering of the output can eliminate this error. Second, the error operator is nearly full rank, while its null space has only logarithmic dimension. Third, the expected error is governed by a single alignment scalar, with a closed-form expression obtained by averaging over random filters. In general, the substitution error asymptotically doubles the output energy, except for filters in the universal zero-error subspace, which incur no error. Collectively, these results show that the substitution error is structured, predictable, and governed by alignment.

URL PDF HTML 收藏
2607.15275 2026-07-17 cs.RO cs.AI cs.LG 新提交

RoboTTT: Context Scaling for Robot Policies

RoboTTT:机器人策略的上下文扩展

Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng, Fengyuan Hu, Yunhao Ge, Jimmy Wu, Tianyuan Dai, Scott Reed, Li Fei-Fei, Yuke Zhu, Linxi "Jim" Fan

机构 * NVIDIA(英伟达公司) Stanford University(斯坦福大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 研究提出RoboTTT,将测试时训练集成到机器人基础模型,通过序列动作强制和截断反向传播扩展视觉运动上下文到8K时间步长,解锁新能力,提升多任务性能,证明上下文长度是机器人基础模型新扩展轴。

Comments Project website: http://research.nvidia.com/labs/gear/robottt/

详情
AI中文摘要

近期的机器人基础模型在单步或短历史视觉运动上下文下运行。我们引入了测试时训练机器人策略(RoboTTT),这是一种机器人模型和训练方法,可将视觉运动上下文扩展到8K时间步长,比现有技术策略高出三个数量级,且不增加推理延迟。在此上下文长度下,解锁了新的机器人能力,如从人类视频演示中一次性上下文模仿、即时策略改进、对扰动的鲁棒性以及在多阶段、长视野任务上更强的性能。还首次观察到随着预训练上下文长度增加,闭环性能稳步提升。核心是将测试时训练集成到机器人基础模型中,通过序列动作强制和截断反向传播来扩展训练上下文长度。在具有挑战性的真实机器人操作任务中,RoboTTT比单步上下文基线的整体性能提高了87%,并完全完成了五分钟、十阶段的装配任务,而基线模型从未做到。用8K时间步长上下文训练的RoboTTT比用1K时间步长预训练的相同模型性能高出62%,表明上下文长度是机器人基础模型新的扩展轴。

英文摘要

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/

URL PDF HTML 收藏
2607.14203 2026-07-17 cs.GR cs.AI cs.CV 新提交

Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation

即时NuRec:用于驾驶场景模拟的前馈3D高斯重建

NVIDIA, :, Jiahui Huang, Jiawei Ren, Michal Tyszkiewicz, Bjoern Haefner, Michael Shelley, Xin Kang, Seung Wook Kim, Ning Xu, Qi Wu, Janick Martinez Esturo, Shengyu Huang, Nick Schneider, Laura Leal-Taixe, Zan Gojcic, Sanja Fidler

机构 * NVIDIA

AI总结 针对自动驾驶3D模拟平台中神经模拟方法存在的速度慢和需逐场景调整问题,提出即时NuRec这一前馈神经重建模型,可快速将多视图驾驶日志转为3DGS世界,在数据集上PSNR表现出色,还能用于闭环模拟。

Comments Project Page: https://research.nvidia.com/labs/sil/projects/instant-nurec/

详情
AI中文摘要

3D模拟平台对自动驾驶至关重要,可实现端到端策略评估,降低开发成本并提高安全性。近年来神经模拟占主导,如NuRec等方法起核心作用,但仍较慢且需逐场景调整。本文提出即时NuRec,一种前馈神经重建模型,能将短多视图驾驶日志在单次前向传播中转换为完全可模拟的3D高斯点云(3DGS)世界。该模型接受校准相机装置的多视图输入,输出包括静态和动态3DGS层、天空立方体贴图及相机ISP校正,通过3DGUT支持非针孔相机模型。它能在约1.5秒内重建10 - 20秒的多相机场景,在Waymo开放数据集上PSNR比最强评估基线高2.01dB。即时NuRec深度集成到NuRec中,与AlpaSim兼容用于闭环模拟。

英文摘要

3D simulation platforms are critical for autonomous driving because they enable end-to-end policy evaluation, thereby reducing development costs and improving safety. In recent years, neural simulation has become predominant, with methods such as NuRec playing a central role; however, these methods remain relatively slow and typically require per-scene tuning. In this work, we present Instant NuRec, a feed-forward neural reconstruction model that turns a short multi-view driving log into a fully simulatable 3D Gaussian Splatting (3DGS) world in a single forward pass. The model accepts multi-view input from a calibrated camera rig and emits a layered output consisting of static and dynamic 3DGS layers, a sky cubemap, and per-camera ISP corrections, while providing native support for non-pinhole camera models via 3DGUT. It reconstructs a 10-20-second multi-camera scene in roughly 1.5 seconds and achieves a PSNR on the Waymo Open Dataset that is 2.01 dB above the strongest evaluated baseline. Instant NuRec is deeply integrated into NuRec and is compatible with AlpaSim for closed-loop simulation.

URL PDF HTML 收藏
2607.12121 2026-07-17 cs.DC cs.LG 版本更新

FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving

FlashDiff:用于扩散模型服务的高效区域执行与调度

Yaqi Qiao, Ping He, Songrun Xie, Ayush Barik, Chensong Zhang, Zhengzhong Tu, Fan Lai

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Vanderbilt University(范德比大学) HKUST(香港科技大学) NVIDIA(英伟达) Texas A&M University(德克萨斯农工大学)

AI总结 研究针对扩散模型服务效率低的问题,提出FlashDiff系统,通过自适应区域执行和调度,利用扩散细化特性及三种机制,有效降低端到端服务延迟,提高吞吐量。

详情
AI中文摘要

扩散模型已成为现代图像、视频和音频生成的核心支柱,但其高效服务仍是挑战。与自回归解码不同,扩散推理在多个去噪步骤中反复更新高维空间或时间潜在变量。现有多GPU并行化方法存在问题。本文提出FlashDiff,通过自适应区域执行和调度提高推理效率。基于扩散细化在潜在区域或去噪步骤中不均匀的观察,FlashDiff利用这些特性选择性执行需进一步细化的区域,并在并发服务请求中重新分配计算空闲时间。它由三种机制组成,包括使用早期注意力信号分解潜在表示、用轻量级运行时控制器估计区域活动、应用亲和感知在线调度器。在实际图像、视频和音频工作负载中,FlashDiff将端到端服务延迟降低30 - 97%,吞吐量提高1.2 - 2.2倍。

英文摘要

Diffusion models have become the central backbone for modern image, video, and audio generation, but their efficient service remains a challenge. Unlike autoregressive decoding, diffusion inference repeatedly updates high-dimensional spatial or temporal latents over many denoising steps. This all-region execution pattern makes generation latency high and limits serving throughput. Existing multi-GPU parallelization methods can reduce per-step computation, but often introduce substantial activation exchange overhead, causing communication to offset or even outweigh the benefits of parallel execution. This paper presents FlashDiff, a diffusion serving system that improves inference efficiency through adaptive regional execution and scheduling. FlashDiff is based on the observation that diffusion refinement is not uniform across latent regions or denoising steps: different regions often stabilize at different rates, while neighboring steps exhibit strong temporal correlation. FlashDiff leverages these properties to selectively execute only regions that require further refinement and to reallocate the resulting compute slack across concurrent serving requests. FlashDiff consists of three mechanisms. First, it decomposes the latent representation into coherent execution regions using early-stage attention signals, preserving semantic structure while exposing fine-grained parallelism. Second, it uses a lightweight runtime controller to estimate region activity and bypass low-impact updates when further refinement is unlikely to affect output quality. Third, it applies an affinity-aware online scheduler that co-locates dependent regions, balances residual load across GPUs, and reuses reclaimed compute capacity to improve serving efficiency. Across real-world image, video, and audio workloads, FlashDiff reduces end-to-end serving latency by 30-97% and improves throughput by 1.2-2.2x.

URL PDF HTML 收藏
2606.29814 2026-07-17 cs.CV 版本更新

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

Nemotron-Labs-Diffusion-Image:推进掩码离散扩散用于高分辨率图像合成

Shufan Li, Greg Heinrich, Hanrong Ye, Yonggan Fu, Aditya Grover, Jan Kautz, Pavlo Molchanov

机构 * NVIDIA

AI总结 提出Nemotron-Labs-Diffusion-Image模型,通过令牌编辑机制和分组交叉熵目标解决掩码离散扩散模型的自校正缺失和训练信号稀疏问题,实现高分辨率文本到图像合成。

Comments 23 pages, 12 figures

详情
AI中文摘要

我们提出Nemotron-Labs-Diffusion-Image,一种用于高分辨率文本到图像合成的最先进的掩码离散扩散模型(MDM)。与先前掩码图像生成工作相比,Nemotron-Labs-Diffusion-Image解决了两个关键挑战。首先,与连续扩散模型在整个图像上逐步细化潜在表示不同,标准MDM缺乏自校正能力,因为离散令牌一旦被取消掩码就无法修改。其次,虽然增加离散图像分词器的词汇量提高了重建保真度,但它给生成建模带来了优化困难,因为每个令牌的训练信号变得越来越稀疏。为了解决第一个挑战,Nemotron-Labs-Diffusion-Image引入了一种令牌编辑机制,使模型能够在推理过程中动态修改已取消掩码的令牌,类似于雕塑家迭代完善其作品。为了解决第二个挑战,我们提出了一种分组交叉熵(GCE)目标,该目标为嵌入空间中接近真实值的令牌分配正学习信号,从而缓解信号稀疏性。为了进一步提高训练效率,我们为GCE实现了一个自定义融合算子,显著减少了大型词汇设置下的VRAM使用。实验结果表明,这些创新显著提高了掩码离散图像生成器的训练效率和图像保真度,在GenEval上达到0.90分,DPG上达到86.9分,HPSv3上达到10.76分。

英文摘要

We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared with prior work on masked image generation, Nemotron-Labs-Diffusion-Image addresses two key challenges. First, unlike continuous diffusion models which progressively refine latent representations across the entire image, standard MDMs lack self-correcting capability because discrete tokens cannot be modified once they are unmasked. Second, although increasing the vocabulary size of discrete image tokenizers improves reconstruction fidelity, it introduces optimization difficulties for generative modeling as the per-token training signal becomes increasingly sparse. To address the first challenge, Nemotron-Labs-Diffusion-Image incorporates a token-editing mechanism that enables the model to dynamically revise already-unmasked tokens during inference, similar to how a sculptor iteratively refines their work. To tackle the second challenge, we propose a Grouped Cross-Entropy (GCE) objective that assigns positive learning signals to tokens neighboring the ground truth in embedding space, thereby alleviating signal sparsity. To further improve training efficiency, we implement a custom fused operator for GCE that significantly reduces VRAM usage in large-vocabulary settings. Experimental results demonstrate that these innovations substantially improve both training efficiency and image fidelity of masked discrete image generators, achieving a score of 0.90 on GenEval, 86.9 on DPG and 10.76 of HPSv3.

URL PDF HTML 收藏
2508.18242 2026-07-17 cs.CV 版本更新

GSVisLoc: Generalizable Visual Localization for Gaussian Splatting Scene Representations

GSVisLoc:用于高斯喷溅场景表示的通用视觉定位

Fadi Khatib, Dror Moran, Guy Trostianetsky, Yoni Kasten, Meirav Galun, Ronen Basri

机构 * Weizmann Institute of Science(魏茨曼科学研究所) NVIDIA(英伟达)

AI总结 GSVisLoc是用于3D高斯喷溅场景表示的视觉定位方法,通过匹配3D高斯生成的场景特征与图像特征来估计相机位姿,分三步进行,无需修改、再训练或额外图像,在标准基准上性能优且能有效推广到新场景。

Comments Accepted to ICCV 2025 Workshops (CALIPOSE). Project page: https://gsvisloc.github.io/

详情
AI中文摘要

我们介绍了GSVisLoc,一种为3D高斯喷溅(3DGS)场景表示设计的视觉定位方法。给定场景的3DGS模型和查询图像,目标是估计相机的位置和方向。通过将场景特征与图像特征稳健匹配来实现,场景特征由3D高斯下采样和编码生成,图像特征通过编码图像块获得。算法分三步:粗匹配、细匹配,最后姿态优化。该方法利用显式3DGS场景表示进行视觉定位,无需修改、再训练或额外参考图像。在室内外场景评估中,在标准基准上有竞争力,优于现有基于3DGS的基线,且能有效推广到新场景。

英文摘要

We introduce GSVisLoc, a visual localization method designed for 3D Gaussian Splatting (3DGS) scene representations. Given a 3DGS model of a scene and a query image, our goal is to estimate the camera's position and orientation. We accomplish this by robustly matching scene features to image features. Scene features are produced by downsampling and encoding the 3D Gaussians while image features are obtained by encoding image patches. Our algorithm proceeds in three steps, starting with coarse matching, then fine matching, and finally by applying pose refinement for an accurate final estimate. Importantly, our method leverages the explicit 3DGS scene representation for visual localization without requiring modifications, retraining, or additional reference images. We evaluate GSVisLoc on both indoor and outdoor scenes, demonstrating competitive localization performance on standard benchmarks while outperforming existing 3DGS-based baselines. Moreover, our approach generalizes effectively to novel scenes without additional training.

URL PDF HTML 收藏
2607.13681 2026-07-16 cs.CV 新提交

Towards Spatial Supersensing in the Wild

迈向野外空间超感知

Tianjun Gu, Tianyu Xin, Kuan Zhang, Bowen Yang, Kok-Chung Chua, Peize Li, Xinran Zhang, Yupeng Chen, Qiyue Zhao, Qinlei Xie, Jianhang Liu, Yucheng Lu, Yinan Han, Marco Pavone, Yiming Li

机构 * Tsinghua University(清华大学) NVIDIA(英伟达) Stanford University(斯坦福大学)

AI总结 研究针对空间超感知中多模态模型基准测试局限于合成视频和家庭场景的问题,引入VSI-Super-Wild基准,受人类认知启发探究世界状态三元组,通过大量真实视频问答对测试发现模型不足及失败模式,为空间超感知发展指明方向。

Comments Accepted to ECCV 2026. Project page: https://vsi-super-wild.github.io/

详情
AI中文摘要

人类能够有效地解析从数小时到数年的连续感官流,构建一个基于空间推理和预测的内部世界模型。为模仿这种能力,空间超感知挑战多模态模型超越语言理解,实现真正的世界建模。然而,其基准测试依赖合成长视频,多限于家庭场景,对现实世界的连续性和多样性探索不足。为此,我们引入VSI-Super-Wild,一个用于评估野外不同场景中长时间空间超感知的大规模基准。受人类构建经验的认知研究启发,我们系统探究世界状态的三元组:智能体、物体和环境。VSI-Super-Wild包含6980个人工验证的问答对,源自442个跨越8个场景类别的真实世界视频。结果显示,尽管静态图像理解有进展,但模型在需要连贯跟踪世界状态随时间变化的任务上持续失败。我们刻画了性能如何随世界状态复杂性和时间跨度下降,并诊断出四种失败模式。这种分类揭示模型缺乏将物体、智能体和环境绑定成统一空间世界模型的机制,这一根本差距为空间超感知指明了前进方向。

英文摘要

Humans can efficiently parse continuous sensory streams, from hours to years, scaffolding an internal world model that grounds spatial reasoning and prediction. To mimic this capacity, spatial supersensing challenges multimodal models to move beyond linguistic understanding toward true world modeling. However, their benchmark relies on synthetic long videos, formed by concatenating random short clips, and is mostly limited to household scenes, leaving real-world continuity and diversity underexplored. To address the gap, we introduce $\textbf{VSI-Super-Wild}$, a large-scale benchmark for evaluating spatial supersensing over long temporal horizons in diverse in-the-wild scenes. Notably, inspired by cognitive studies on how humans structure experience, we systematically probe the full triad of world state: the agent (observer), objects (scene items), and the environment (places and global layout). In total, VSI-Super-Wild contains $\textbf{6,980}$ human-verified question-answer pairs derived from $\textbf{442}$ real-world videos spanning 8 scene categories, including long-form recordings exceeding 4 hours. Results on VSI-Super-Wild expose a fundamental disconnect: despite advances in static image understanding, models consistently fail at tasks that require coherent world-state tracking over time. We characterize how performance degrades with world-state complexity and temporal horizon, and diagnose four failure modes: spatial collapse, semantic shortcuts, insufficient update, and instance confusion. This taxonomy reveals that models lack mechanisms to bind objects, agents, and environments into a unified spatial world model, a fundamental gap that defines the path forward for spatial supersensing.

URL PDF HTML 收藏
2607.06405 2026-07-16 cs.MM cs.SD 版本更新

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space

在潜在空间中通过跨模态对齐实现精确的视频到音频生成

Thanh V. T. Tran, Ngoc-Son Nguyen, Luong Tran, Long-Khanh Pham, Paarth Neekhara, Shehzeen Hussain, Van Nguyen

机构 * FPT Software AI Center(FPT软件人工智能中心) NVIDIA Corporation(NVIDIA公司)

AI总结 研究视频到音频生成问题,提出Flowley架构,结合视觉特征与文本提示,通过渐进软掩码交叉注意力实现视听同步,无额外计算成本,还提出SoundCap字幕,该方法在多个指标及零样本音频质量上达先进水平。

Comments Accepted to ECCV 2026

详情
AI中文摘要

视频到音频(V2A)生成旨在合成与无声视频语义一致且时间同步的逼真音频。尽管有进展,但许多方法仍存在多阶段训练成本高、运行时间长或牺牲细粒度时间线索等问题。为此提出Flowley,一种端到端单阶段训练架构,结合视觉特征和文本提示生成音轨。引入渐进软掩码交叉注意力,直接在注意力机制中嵌入视听同步,无额外计算成本。还指出现有V2A基准缺乏声音描述性字幕,提出SoundCap创建详细的声音感知字幕。Flowley在多个指标上实现了最先进性能,结合SoundCap在零样本设置下音频质量超越现有方法。

英文摘要

Video-to-audio (V2A) generation aims to synthesize realistic audio that is both semantically consistent with and temporally synchronized to a silent video. Despite recent progress, many methods still rely on multi-stage training, resulting in high computational costs and long runtimes, or transform visual input into text to leverage pretrained text-to-audio models, sacrificing fine-grained temporal cues. To overcome these limitations, we propose Flowley, an end-to-end, single-stage training architecture that produces soundtracks by combining visual features with textual prompts. Crucially, we introduce Progressive Soft-masked Cross-Attention, which embeds audio-visual synchronization directly within its attention mechanism, adding zero additional computational cost compared to standard attention layers. We further observe that existing V2A benchmarks lack sound-oriented descriptive captions, which can potentially degrade the quality of the synthesized audio. To remedy this, we propose SoundCap, a plug-and-play pipeline for creating detailed, sound-aware captions that guide the model. Remarkably, without integrating any pretrained audio-visual alignment modules, Flowley achieves state-of-the-art performance on VGGSound across multiple metrics. Moreover, by incorporating SoundCap, we further exceed the performance of the strongest existing close-sourced methods in terms of audio quality in the zero-shot setting.

URL PDF HTML 收藏
2606.31043 2026-07-16 cs.LG cs.RO 版本更新

Warp RL: Reshaping Base Policy Distributions for Dynamics Adaptation

Warp RL: 重塑基础策略分布以进行动力学自适应

Ethan Hirschowitz, Fabio Ramos

机构 * University of Sydney(悉尼大学) NVIDIA(英伟达)

AI总结 针对残差强化学习在动力学偏移下无法调整分布形状的问题,提出Warp RL方法,通过可逆状态条件变换重塑基础策略的动作分布,在ManiSkill3任务和真实机器人插销任务中优于残差校正。

Comments 17 pages, 7 figures

详情
AI中文摘要

残差强化学习通过学习对其动作的加性校正来适应预训练的机器人策略。当适应相当于移动基础策略的动作分布时,加性校正是有效的,但加性校正无法改变分布的形状、尺度或状态依赖的几何结构——我们将这些局限性形式化为错误的方差、不正确的置信度和非均匀校正。我们证明这些在动力学偏移下很重要:当基础分布在几何上与偏移系统不匹配时,残差校正甚至可能不如未适应的策略。我们提出\textbf{Warp RL},一种策略适应方法,用基础策略动作分布的可逆、状态条件变换替代加性残差。通过单调有理二次样条流[ arXiv:0706.1234v1 ]实例化,Warp RL保持恒等初始化,严格推广加性残差校正,并暴露了一个适用于策略梯度和无梯度优化的结构化适应空间。在具有受控动力学偏移的各种ManiSkill3操作任务中,当平移足够时,Warp RL匹配残差校正,而当适应需要分布重塑时,其性能显著优于残差校正。我们进一步证明,在离策略的仿真到现实流程中,变形可以替代加性校正,在真实机器人插销任务中实现相当的成功率,同时任务完成速度提高30%。

英文摘要

Residual reinforcement learning adapts a pretrained robot policy by learning an additive correction to its actions. While effective when adaptation amounts to shifting the base policy's action distribution, additive corrections cannot change the distribution's shape, scale, or state-dependent geometry -- limitations we formalize as wrong variance, miscalibrated confidence, and non-uniform correction. We show that these matter under dynamics shift: when the base distribution is geometrically mismatched to the shifted system, residual correction can underperform even the unadapted policy. We propose Warp RL, a policy adaptation method that replaces additive residuals with an invertible, state-conditioned transformation of the base policy's action distribution. Instantiated with monotonic rational-quadratic spline flows (arXiv:1906.04032), Warp RL preserves identity initialization, strictly generalizes additive residual correction, and exposes a structured adaptation space suitable for both policy-gradient and gradient-free optimization. Across a variety of ManiSkill3 manipulation tasks with controlled dynamics shifts, Warp RL matches residual correction when translation is sufficient and substantially outperforms it when adaptation requires distributional reshaping. We further demonstrate that warping can replace additive correction in an off-policy sim-to-real pipeline, achieving comparable success rate with 30% faster task completion on a real-robot peg-insertion task.

URL PDF HTML 收藏
2211.14939 2026-07-15 cs.LG q-bio.BM

Applying Deep Reinforcement Learning to the HP Model for Protein Structure Prediction

将深度强化学习应用于HP模型进行蛋白质结构预测

Kaiyuan Yang, Houjing Huang, Olafs Vandans, Adithya Murali, Fujia Tian, Roland H. C. Yap, Liang Dai

机构 * Department of Computer Science, School of Computing, National University of Singapore(新加坡国立大学计算机科学系) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) NVIDIA Seattle Robotics Lab(NVIDIA西雅图机器人实验室) Department of Physics, City University of Hong Kong(香港城市大学物理系)

AI总结 本研究利用深度强化学习解决HP模型中的蛋白质结构预测问题,通过深度Q网络和LSTM架构提升搜索效率,找到多个最佳解。

Comments Published at Physica A: Statistical Mechanics and its Applications, available online 7 December 2022. Extended abstract accepted by the Machine Learning and the Physical Sciences workshop, NeurIPS 2022

Journal ref Physica A: Statistical Mechanics and its Applications 609 (2023) 128395

详情
AI中文摘要

计算生物物理领域的一个核心问题是蛋白质结构预测,即找到给定氨基酸序列的最优折叠方式。这个问题在经典抽象模型HP模型中已被研究,其中蛋白质被建模为在晶格上的H(疏水性)和P(极性)氨基酸序列。目标是找到最大化H-H接触的构型。已知即使在这种简化设定中,该问题也是不可解的(NP难)。在本工作中,我们应用深度强化学习(DRL)到二维HP模型。我们能够获得已知最佳能量的HP序列构型,这些序列长度从20到50不等。我们的DRL基于深度Q网络(DQN)。我们发现基于长短期记忆(LSTM)架构的DQN显著增强了RL学习能力,并显著提高了搜索过程。DRL可以高效地采样状态空间,而无需手动启发式方法。实验表明,它可以在每次试验中找到多个不同的最佳已知解。本研究展示了深度强化学习在蛋白质折叠的HP模型中的有效性。

英文摘要

A central problem in computational biophysics is protein structure prediction, i.e., finding the optimal folding of a given amino acid sequence. This problem has been studied in a classical abstract model, the HP model, where the protein is modeled as a sequence of H (hydrophobic) and P (polar) amino acids on a lattice. The objective is to find conformations maximizing H-H contacts. It is known that even in this reduced setting, the problem is intractable (NP-hard). In this work, we apply deep reinforcement learning (DRL) to the two-dimensional HP model. We can obtain the conformations of best known energies for benchmark HP sequences with lengths from 20 to 50. Our DRL is based on a deep Q-network (DQN). We find that a DQN based on long short-term memory (LSTM) architecture greatly enhances the RL learning ability and significantly improves the search process. DRL can sample the state space efficiently, without the need of manual heuristics. Experimentally we show that it can find multiple distinct best-known solutions per trial. This study demonstrates the effectiveness of deep reinforcement learning in the HP model for protein folding.

URL PDF HTML 收藏
2607.11533 2026-07-14 cs.CV cs.LG 新提交

Adaptive Routing for Efficient Diffusion Transformer-Based PNI Prediction

基于高效扩散变压器的PNI预测的自适应路由

Youngung Han, Dohyun Kweon, Kyeonghun Kim, Hyunsu Go, Jina Jeong, Suah Park, Induk Um, Junga Kim, Anna Jung, Yului Jeong, Sungha Park, Jinyong Jun, Pa Hong, Woo Kyoung Jeong, Won Jae Lee, Ken Ying-Kai Liao, Hyuk-Jae Lee, Nam-Joon Kim

机构 * Seoul National University(首尔国立大学) Kyung Hee University(庆熙大学) OUTTA Chung-Ang University(Chung-Ang 大学) Seoul National University School of Medicine(首尔国立大学医学院) Samsung Changwon Hospital(三星昌原医院) Samsung Medical Center(三星医疗中心) NVIDIA AI Technology Center(NVIDIA AI 技术中心)

AI总结 针对胆管癌PNI术前MRI预测难题,传统方法有局限。本文将PNI预测设为扩散分类问题,用基于Transformer的表示实现去噪网络,并引入自适应路由提高效率,实验取得了0.731的AUC及257.57 GFLOPs的结果。

详情
AI中文摘要

神经周围侵犯(PNI)是胆管癌的关键预后因素。然而,由于细微的成像特征超出肿瘤边界延伸到周围区域,从磁共振成像(MRI)进行术前预测仍然具有挑战性。传统卷积神经网络在捕捉远距离空间依赖性方面有限。基于Transformer的架构通过聚合空间分布的上下文线索改进了体积MRI的全局建模,但在肿瘤周围区域捕捉细微和噪声敏感模式仍具挑战。基于扩散的分类器通过利用基于去噪的类评分提供了一种替代方案。但这些方法由于基于Transformer的建模和迭代去噪过程的结合而引入了大量计算开销。为应对这些挑战,我们将PNI预测公式化为基于扩散的分类问题,并使用基于Transformer的表示实现去噪网络。为提高计算效率,我们引入了跨注意力头、空间令牌和MLP宽度的自适应路由。实验结果表明,该方法在257.57 GFLOPs下实现了0.731的AUC。

英文摘要

Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. However, its preoperative prediction from magnetic resonance imaging (MRI) remains challenging due to subtle imaging features that extend beyond tumor boundaries into surrounding regions. Conventional convolutional neural networks are limited in capturing long-range spatial dependencies. Transformer-based architectures improve global modeling of volumetric MRI by aggregating spatially distributed contextual cues, yet capturing subtle and noise-sensitive patterns in peritumoral regions remains challenging. Diffusion-based classifiers offer an alternative formulation by leveraging denoising-based class scoring to better capture such subtle patterns. However, these approaches introduce substantial computational overhead due to the combination of transformer-based modeling and iterative denoising processes. To address these challenges, we formulate PNI prediction as a diffusion-based classification problem and implement the denoising network using a transformer-based representation. To improve computational efficiency, we introduce adaptive routing across attention heads, spatial tokens, and MLP width. Experimental results demonstrate that the proposed approach achieves an AUC of 0.731 with 257.57 GFLOPs.

URL PDF HTML 收藏
2607.10992 2026-07-14 cs.CV cs.AI 新提交

LoSA-Net: A Localized and Scale-Adaptive Network for Boundary-Sensitive Prediction of Perineural Invasion in 3D MRI

LoSA-Net:用于3D MRI中神经周围侵犯边界敏感预测的局部化和尺度自适应网络

Youngung Han, Hyunsu Go, Kyeonghun Kim, Induk Um, Junga Kim, Jaewon Jung, Woo Kyoung Jeong, Won Jae Lee, Pa Hong, Ken Ying-Kai Liao, Hyuk-Jae Lee, Nam-Joon Kim

机构 * Seoul National University(首尔国立大学) OUTTA Chung-Ang University(Chung-Ang 大学) Samsung Medical Center, Sungkyunkwan University School of Medicine(三星医疗中心,成均馆大学医学院) Samsung Changwon Hospital(三星昌原医院) NVIDIA AI Technology Center(NVIDIA AI 技术中心)

AI总结 研究针对3D MRI中神经周围侵犯(PNI)预测难题,提出LoSA-Net架构,通过TNA、SAFM和CSRA技术,在168例胆管癌患者的对比增强MRI扫描中,该网络AUC达0.7567,优于卷积和Transformer基线。

Comments Published in the 2026 IEEE 23rd International Symposium on Biomedical Imaging (ISBI 2026); accepted for oral presentation

Journal ref 2026 IEEE 23rd International Symposium on Biomedical Imaging (ISBI), 2026

详情
AI中文摘要

神经周围侵犯(PNI)是肿瘤侵袭性的临床相关指标,影响手术决策,因此可靠的术前评估很重要。然而,PNI在MRI上的细微特征常与附近解剖结构相似,常规下采样或过度全局特征聚合会削弱这些精细神经周围线索,降低传统体积模型的有效性。我们提出LoSA-Net,一种用于3D MRI中边界敏感PNI预测的局部化和尺度自适应架构。其中,Talking Neighborhood Attention(TNA)通过局部自注意力和逐头混合保留神经对齐细节,Scale-Adaptive Feature Mixing(SAFM)使用多尺度深度处理调节感受野,Cross-Scale Refinement and Alignment(CSRA)在各阶段保持语义上下文和高分辨率边界之间的一致性。在168例胆管癌患者的对比增强MRI扫描中,LoSA-Net的AUC为0.7567,在匹配的预处理和优化设置下优于代表性的卷积和Transformer基线。

英文摘要

Perineural invasion (PNI) is a clinically relevant indicator of tumor aggressiveness and can influence surgical decision-making, motivating interest in reliable preoperative assessment. The subtle MRI features of PNI, however, often resemble nearby anatomy, complicating noninvasive prediction. These fine perineural cues are easily attenuated by routine downsampling or overly global feature aggregation, reducing the effectiveness of conventional volumetric models. We present LoSA-Net, a localized and scale-adaptive architecture for boundary-sensitive PNI prediction in 3D MRI. Talking Neighborhood Attention (TNA) preserves nerve-aligned detail through localized self-attention with head-wise mixing, and Scale-Adaptive Feature Mixing (SAFM) modulates the receptive field using multi-scale depthwise processing. Cross-Scale Refinement and Alignment (CSRA) maintains consistency between semantic context and high-resolution boundaries across stages. In contrast-enhanced MRI scans from 168 patients with cholangiocarcinoma, LoSA-Net achieves an AUC of 0.7567 and outperforms representative convolutional and transformer baselines under matched preprocessing and optimization settings.

URL PDF HTML 收藏
2607.10988 2026-07-14 cs.CV cs.AI 新提交

MMA-Former: Multi-Window Mixture-of-Head Attention Transformer for Adaptive PNI Prediction in 3D MRI

MMA-Former:用于3D MRI中自适应PNI预测的多窗口混合注意力头变压器

Youngung Han, Induk Um, Kyeonghun Kim, Junga Kim, Hyunsu Go, Jaewon Jung, Woo Kyoung Jeong, Won Jae Lee, Pa Hong, Ken Ying-Kai Liao, Hyuk-Jae Lee, Nam-Joon Kim

机构 * Seoul National University(首尔国立大学) OUTTA Chung-Ang University(Chung-Ang 大学) Samsung Medical Center, Sungkyunkwan University School of Medicine(三星医疗中心,全北大学医学院) Samsung Changwon Hospital(三星昌原医院) NVIDIA AI Technology Center(NVIDIA AI 技术中心)

AI总结 针对3D MRI无创预测PNI的挑战,提出MMA-Former,采用粗-细变压器结构及窗口特定混合注意力机制,能并行多尺度提取特征,实现空间自适应特征提取,在回顾性数据集上AUC达0.752,优于其他架构。

Comments Published in the 2026 IEEE 23rd International Symposium on Biomedical Imaging (ISBI 2026); accepted for oral presentation

Journal ref 2026 IEEE 23rd International Symposium on Biomedical Imaging (ISBI), 2026

详情
AI中文摘要

神经周围侵犯(PNI)是胆管癌的关键预后因素。从3D MRI进行无创预测具有挑战性,需要能有效捕捉细粒度细节和全局上下文的模型。我们提出了多窗口混合注意力头变压器(MMA-Former),这是一种新颖的端到端3D架构,具有用于并行多尺度特征提取的粗-细变压器(CFT)结构。我们通过集成一种新颖的窗口特定混合注意力(WS-MoH)机制来改进此结构。与标准多头自注意力(MSA)不同,WS-MoH为每个3D窗口生成一个表示,并将整个窗口动态路由到专门的或通用的注意力头。这实现了针对每个窗口的局部上下文进行空间自适应特征提取,在不增加参数的情况下增强了专业化并减少了冗余。在168例T1加权MRI扫描的回顾性数据集中进行评估,MMA-Former的AUC为0.752,优于其他3D架构,包括最佳的CNN(AUC为0.708)和变压器基线(AUC为0.681)。

英文摘要

Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. Non-invasive prediction from 3D MRI is challenging, demanding models that efficiently capture both fine-grained details and global context. We propose the Multi-window Mixture-of-Head Attention Transformer (MMA-Former), a novel end-to-end 3D architecture featuring a Coarse-Fine Transformer (CFT) structure for parallel multi-scale feature extraction. We advance this structure by integrating a novel Window-Specific Mixture-of-Head attention (WS-MoH) mechanism. Unlike standard Multi-Head Self Attention (MSA), WS-MoH generates a representation for each 3D window and dynamically routes the entire window to specialized or common attention heads. This enables spatially adaptive feature extraction tailored to the local context of each window, enhancing specialization and reducing redundancy without increasing parameters. Evaluated on a retrospective dataset of 168 T1-weighted MRI scans, MMA-Former achieved an AUC of 0.752, outperforming other 3D architectures, including the best CNN (AUC of 0.708) and Transformer baselines (AUC of 0.681).

URL PDF HTML 收藏
2607.10044 2026-07-14 cs.LG 新提交

FlashTrie: A GPU-Accelerated Constrained Beam Search for Generative Retrieval

FlashTrie:用于生成式检索的GPU加速约束束搜索

Dakshitha Anandakumar, Anurag Mukkara, Wenxiang Hu, Jiusheng Chen, M Akash Kumar, Ting Ye, Qiang Lou, Jian Jiao

机构 * Microsoft(微软) Nvidia(英伟达)

AI总结 研究针对生成式检索中约束解码的瓶颈问题,提出FlashTrie方法。通过优化GPU上的约束束搜索,采用整数感知简洁trie布局和协作CUDA内核等技术,显著降低解码延迟、提高吞吐量,在实验中取得良好效果并提升了收入。

详情
AI中文摘要

约束解码在生成式检索中至关重要,直接从查询生成的文档标识符必须与预定义的有效ID库完全匹配。大规模时,通常使用带有束搜索的trie进行约束解码,但大多数实现运行在CPU上。随着束宽度增加,有限的并行性使trie遍历和候选验证成为服务瓶颈。我们提出FlashTrie,通过在GPU上优化约束束搜索来解决此限制。它引入整数感知简洁trie布局,使用位压缩减少内存占用,同时将完整索引保存在GPU高带宽内存中以减少内存停顿;还引入协作CUDA内核,完全在设备上执行束扩展、验证和修剪,无需主机逐步骤编排。它进一步用GPU感知并行原语取代CPU风格的不规则查找和堆维护,提高线程利用率并减少分歧。这些设计显著降低解码延迟并提高吞吐量,同时保持检索质量。在包含8亿关键词且束宽度高达1000的库上,FlashTrie将trie搜索延迟降低到3毫秒以下,比高度优化的多线程CPU基线实现高达24倍的加速。这些改进使FlashTrie在诸如赞助搜索等延迟关键应用中能够将束大小扩展多达5倍。在一个流行商业搜索引擎上的大规模在线A/B实验中,它带来了统计学上显著的0.71%的收入提升,实现了以前仅离线可行规模的实时约束解码。FlashTrie代码将在评审过程后公开发布。

英文摘要

Constrained decoding is essential in generative retrieval, where document identifiers generated directly from a query must exactly match a predefined library of valid IDs. At scale, decoding is often constrained using a trie with beam search but most implementations run on CPU. Limited parallelism then makes trie traversal and candidate validation a serving bottleneck as beam width grows. We present FlashTrie, which addresses this limitation by optimizing constrained beam search on GPUs. It introduces an integer-aware succinct trie layout that uses bit compression to reduce memory footprint while keeping the full index in GPU high-bandwidth memory reducing memory stalls, and a cooperative CUDA kernel that performs beam expansion, validation, and pruning entirely on-device without per-step host orchestration. It further replaces CPU-style irregular lookup and heap maintenance with GPU-aware parallel primitives, improving warp utilization and reducing divergence. Together, these designs significantly reduce decoding latency and increase throughput while preserving retrieval quality. On a library of 800M keywords with beam widths up to 1000, FlashTrie reduces trie-search latency to under 3 ms, achieving up to 24x speedup over a highly optimized multi-threaded CPU baseline. These improvements enable FlashTrie to scale beam sizes by up to 5x in latency-critical applications such as sponsored search. In a large-scale online A/B experiment on a popular commercial search engine, it delivers a statistically significant +0.71% revenue lift, enabling real-time constrained decoding at a scale previously feasible only offline. The FlashTrie code will be publicly released after the review process.

URL PDF HTML 收藏
2607.09739 2026-07-14 cs.AI cs.CL 新提交

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

分数集之前的核心集:用于大语言模型基准测试的评估无监督提示子集选择

Jihan Yao, Gantavya Bhatt, Arnav Das, Peter Jin, Ke Bao, Qiaolin Yu, Khushi Bhardwaj, Chang Su, Jialei Wang, Yikai Zhu, Sugam Devare, Damon Mosk-Aoyama, Zhen Dong, Venkat Krishna Srinivasan, Yineng Zhang, Oleksii Kuchaiev, Jiantao Jiao, Banghua Zhu, Jeff Bilmes

机构 * University of Washington(华盛顿大学) University of California, Berkeley(加利福尼亚大学伯克利分校) Oracle(甲骨文公司) Together AI(Together AI公司) LMSYS NVIDIA(英伟达公司)

AI总结 研究大语言模型基准测试的核心集选择,采用评估无监督方法,利用次模子集选择,开发多种次模函数。在新大规模套件上实验发现设施选址函数效果好,该目标不限于特定模式,在相关排行榜上表现优且计算成本低,证明次模性对基准压缩有用。

详情
AI中文摘要

我们研究大语言模型基准核心集选择问题,即在多个基准测试中选择一小部分提示,使诱导的模型分数和排名接近完整基准测试集的结果。在评估无监督基准核心集选择中,选择算法不使用模型评估结果,通过在多个基准测试中生成提示子集进行细粒度操作。我们使用次模子集选择,并为此开发和评估了许多不同的次模函数。在一个包含35个异构基准测试、18个前沿大语言模型和超61K提示的新大规模套件上,我们发现仅基于廉价语义提示嵌入操作的设施选址函数在一系列核心集预算下比12个基于分数和多样性的基线更好地保留大语言模型分数。此外,我们提出的目标不限于评估无监督模式,在仅需选择少数完整基准测试且有大量模型分数可用的设置中,相同目标在MMLU和MTEB排行榜上与现有最佳基线相当或更优,且计算成本更低。我们的结果表明,一般来说,次模性是基准压缩的强大且可靠工具。

英文摘要

We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite. In evaluation-unsupervised benchmark coreset selection (our approach), the selection algorithm uses no model evaluation outcomes, and operates on a fine granularity by producing subsets of prompts over multiple benchmarks rather than producing a sub-collection of entire benchmarks. We use submodular subset selection, and we develop and evaluate many different submodular functions for this purpose, including determinantal point process (DPP) based approaches, submodular mutual information functions, and facility location-based functions. On a new large-scale suite of 35 heterogeneous benchmarks spanning five different capability categories, 18 frontier LLMs, and over 61K prompts, we find that the facility location (FL) function operating exclusively on inexpensive semantic prompt embeddings preserves LLM scores better than twelve separate score-based and diversity-based baselines, across a range of coreset budgets. Moreover, we show our proposed objective is not limited to the evaluation-unsupervised regime: in the setting where only a handful of whole benchmarks must be selected and a large amount of model scores are available, the same objective matches or outperforms state-of-the-art baselines on the MMLU and MTEB leaderboards, while being substantially cheaper to compute. Together, our results suggest that submodularity, in general, is a strong and reliable tool for benchmark compression.

URL PDF HTML 收藏
2605.18601 2026-07-14 cs.CV 版本更新

Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models

Incantation: 自然语言作为多实体视频世界模型的动作接口

Shangwen Zhu, Qianyu Peng, Zhao Pu, Zhilei Shu, Xiangrui Ke, Zhaohu Xing, Zizhao Tong, Zeqing Wang, Xinyu Cui, Zian Zheng, Huangji Wang, Jian Zhao, Yeying Jin, Fan Cheng, Ruili Feng

机构 * SJTU(上海交通大学) NVIDIA Research(英伟达研究) USTC(中国科学技术大学) UCAS(乌兹别克斯坦科学院) NUS(新加坡国立大学) UWaterloo(滑铁卢大学) HKUST(香港理工大学) HKU(香港大学) ZGCA(浙江大学)

AI总结 本研究提出了一种基于自然语言的动作接口,用于多实体视频世界模型,解决了传统接口在细粒度多实体控制和跨实体、跨世界泛化能力上的不足,通过引入自然语言条件化实现了更强大的表达能力。

详情
AI中文摘要

现代交互式视频世界模型已实现了令人印象深刻的视觉保真度,但缺乏细粒度的多实体控制和跨实体、跨世界的泛化能力。我们追溯这一差距到动作接口:标准控制协议(例如动画ID、设备输入、场景级标题)在设计时将动作语义绑定到特定实体或引擎。我们提出自然语言作为接口,以解锁任何先前接口都无法实现的表达能力,并展示了Incantation,第一个具有每潜在帧(0.25秒)自然语言条件化的交互式视频世界模型,支持同时多实体控制和概念级跨实体转移,超越任何固定的渲染管道。我们配对了一个预训练的双向视频主干与帧本地文本交叉注意力,并通过ODE初始化的Self-Forcing蒸馏与RoPE解耦的滑动KV缓存实现实时长时间跨度流媒体。我们在跨实体转移(89% vs. 43%)和out-of-vocabulary提示(90% vs. 0%)上超越了Action-Index基线,并且我们的两步学生在480p下以19.7 FPS稳定运行,FVD在2小时滚动中保持稳定。我们进一步将相同的架构和训练配方应用于《国王之剑》,仅更改每个实体的动作词汇槽。我们已发布Incantation数据集的预览子集,包含手动收集的《艾尔登法环》玩家-Boss战斗片段,带有结构化的动作导向元数据。更大规模的《艾尔登法环》和KOF数据将在完整项目中发布。

英文摘要

Modern interactive video world models have achieved impressive visual fidelity, yet lack fine-grained multi-entity control and cross-entity, cross-world generalization. We trace this gap to the action interface: standard control protocols (e.g. animation IDs, device inputs, scene-level captions) bind action semantics to specific entities or engines at design time. We propose natural language as the interface to unlock expressiveness that no prior interface can achieve, and we present Incantation, the first interactive video world model with per-latent-frame (0.25 s) natural-language conditioning that supports simultaneous multi-entity control and concept-level cross-entity transfer beyond any fixed rendering pipeline. We pair a pretrained bidirectional video backbone with frame-local text cross-attention, and enable real-time long-horizon streaming through ODE-initialized Self-Forcing distillation with a RoPE-decoupled sliding KV-cache. We surpass the Action-Index baseline on cross-entity transfer (89% vs. 43%) and out-of-vocabulary prompts (90% vs. 0%), and our 2-step student sustains 19.7 FPS at 480p with stable FVD over 2-hour rollouts. We further apply the same architecture and training recipe to The King of Fighters, changing only the per-entity action vocabulary slots. We have released a preview subset of the Incantation dataset at https://huggingface.co/datasets/zhush/incantation-elden-ring-scenes, containing manually collected Elden Ring player-boss combat clips with structured action-oriented metadata. Larger-scale Elden Ring and KOF data will be released with the full project.

URL PDF HTML 收藏
2604.25917 2026-07-14 cs.AI cs.CL cs.LG 版本更新

Recursive Multi-Agent Systems

递归多智能体系统

Jiaru Zou, Rui Pan, Ruizhong Qiu, Pan Lu, Shizhe Diao, Jindong Jiang, Hanghang Tong, Tong Zhang, Markus J. Buehler, Jingrui He, James Zou

机构 * UIUC(伊利诺伊大学香槟分校) Stanford University(斯坦福大学) NVIDIA(英伟达) MIT(麻省理工学院)

AI总结 本文提出递归多智能体框架RecursiveMAS,通过递归计算提升多智能体协作效率,实验证明其在多个基准测试中准确率提升8.3%,推理速度提升1.2-2.4倍,token使用减少34.6%-75.6%。

Comments Project Website: https://recursivemas.github.io

详情
AI中文摘要

递归或循环语言模型最近通过迭代优化相同模型计算来加深推理,我们将其扩展到多智能体系统,探讨智能体协作能否通过递归扩展。我们引入RecursiveMAS框架,将整个系统视为统一的潜在空间递归计算。通过轻量级RecursiveLink模块连接异构智能体,实现分布内潜在思维生成和跨智能体潜在状态转移。为优化框架,我们开发了内-外循环学习算法,通过共享梯度分配进行迭代整体优化。理论分析显示RecursiveMAS比标准文本基多智能体系统更高效,递归训练中保持稳定梯度。实验中,我们基于4种代表性智能体协作模式在9个基准测试中评估,相比先进单/多智能体和递归计算基线,RecursiveMAS在准确率、推理速度和token使用上均取得显著提升。代码和数据见https://recursivemas.github.io。

英文摘要

Recursive or looped language models have recently emerged as a new scaling axis by iteratively refining the same model computation over latent states to deepen reasoning. We extend such scaling principle from a single model to multi-agent systems, and ask: Can agent collaboration itself be scaled through recursion? To this end, we introduce RecursiveMAS, a recursive multi-agent framework that casts the entire system as a unified latent-space recursive computation. RecursiveMAS connects heterogeneous agents as a collaboration loop through the lightweight RecursiveLink module, enabling in-distribution latent thoughts generation and cross-agent latent state transfer. To optimize our framework, we develop an inner-outer loop learning algorithm for iterative whole-system co-optimization through shared gradient-based credit assignment across recursion rounds. Theoretical analyses of runtime complexity and learning dynamics establish that RecursiveMAS is more efficient than standard text-based MAS and maintains stable gradients during recursive training. Empirically, we instantiate RecursiveMAS under 4 representative agent collaboration patterns and evaluate across 9 benchmarks spanning mathematics, science, medicine, search, and code generation. In comparison with advanced single/multi-agent and recursive computation baselines, RecursiveMAS consistently delivers an average accuracy improvement of 8.3%, together with 1.2$\times$-2.4$\times$ end-to-end inference speedup, and 34.6%-75.6% token usage reduction. Code and Data are provided in https://recursivemas.github.io.

URL PDF HTML 收藏
2601.14046 2026-07-14 cs.CL cs.SD 版本更新

PRiSM: Benchmarking Phone Realization in Speech Models

PRiSM:语音模型中电话实现的基准测试

Shikhar Bharadwaj, Chin-Jou Li, Yoonjae Kim, Kwanghee Choi, Eunjung Yeo, Ryan Soh-Eun Shim, Hanyu Zhou, Brendon Boldt, Karen Rosero Jacome, Kalvin Chang, Darsh Agrawal, Keer Xu, Chao-Han Huck Yang, Jian Zhu, Shinji Watanabe, David R. Mortensen

机构 * CMU(卡内基梅隆大学) Gwangju Institute of Science and Technology(光州科学技术院) UT Austin(得克萨斯大学奥斯汀分校) LMU Munich(慕尼黑路德维希-马克西米利安大学) UC Berkeley(伯克利加州大学) NVIDIA(英伟达) UBC(不列颠哥伦比亚大学)

AI总结 该研究针对语音模型中电话实现进行基准测试,引入PRiSM开源基准用内在和外在评估揭示语音感知盲点,标准化评估并通过探针评估下游效用,发现训练语言接触对PR性能关键,编码器-CTC模型稳定,专门PR模型优于大型音频语言模型,还发布相关资源推动多语言语音模型发展。

Comments Presented at ACL 2026

详情
AI中文摘要

电话识别(PR)是跨语言语音处理和语音分析中与语言无关建模的原子接口。尽管在开发PR系统方面付出了长期努力,但目前的评估仅衡量表面转录准确性。我们引入了PRiSM,这是第一个开源基准,旨在通过对PR系统的内在和外在评估来揭示语音感知中的盲点。PRiSM标准化了基于转录的评估,并通过转录和表示探针评估临床、教育和多语言环境中的下游效用。我们发现训练期间的多样化语言接触是PR性能的关键,编码器-CTC模型最稳定,专门的PR模型仍优于大型音频语言模型。PRiSM发布代码、配方和数据集,以推动该领域朝着具有强大语音能力的多语言语音模型发展。

英文摘要

Phone recognition (PR) serves as the atomic interface for language-agnostic modeling for cross-lingual speech processing and phonetic analysis. Despite prolonged efforts in developing PR systems, current evaluations only measure surface-level transcription accuracy. We introduce PRiSM, the first open-source benchmark designed to expose blind spots in phonetic perception through intrinsic and extrinsic evaluation of PR systems. PRiSM standardizes transcription-based evaluation and assesses downstream utility in clinical, educational, and multilingual settings with transcription and representation probes. We find that diverse language exposure during training is key to PR performance, encoder-CTC models are the most stable, and specialized PR models still outperform Large Audio Language Models. PRiSM releases code, recipes, and datasets to move the field toward multilingual speech models with robust phonetic ability: https://github.com/changelinglab/prism.

URL PDF HTML 收藏
2601.13132 2026-07-14 cs.CV 版本更新

SplatReasoner: Enhancing Embodied Reasoning and Grounding by Novel View Synthesis

SplatReasoner:通过新颖视图合成增强具身推理与基础能力

Kim Yu-Ji, Dahye Lee, Kim Jun-Seong, Nam Hyeon-Woo, GeonU Kim, Yongjin Kwon, Yu-Chiang Frank Wang, Jaesung Choe, Tae-Hyun Oh

机构 * POSTECH KAIST(韩国科学技术院) ETRI(韩国电子电信研究院) NVIDIA(英伟达)

AI总结 研究针对视觉语言模型应用于具身场景理解受固定视角限制的问题,提出SplatReasoner框架,利用3D高斯点云将新颖视图合成引入推理过程,经实验验证该方法能提升具身推理和3D基础能力。

Comments Accepted at ECCV 2026. Project page: https://splatreasoner.github.io/

详情
AI中文摘要

视觉语言模型(VLMs)在图像和视频上展现出强大推理能力,但应用于具身场景理解时,常受限于情景RGB-D记忆中的固定视角。这些观察可能因遮挡、物体截断、视野受限或视图组合不佳而无法捕捉与查询相关的证据。我们提出SplatReasoner框架,通过利用3D高斯点云(3DGS)将新颖视图合成引入VLM推理过程。给定关于3D场景的用户查询,SplatReasoner检索相关观察并合成查询条件视角,以揭示回答查询和在3D中定位所指实体所需的视觉证据。实验表明,查询条件新颖视图合成在固定视角记忆和语言嵌入3DGS基线之上,提升了具身推理和3D基础能力。

英文摘要

Vision-Language Models (VLMs) have demonstrated strong reasoning capabilities over images and videos, yet their application to embodied scene understanding often constrained by the fixed viewpoints stored in episodic RGB-D memories. These observations may fail to capture query-relevant evidence due to occlusions, object truncation, restricted fields of view, or suboptimal view composition. We present SplatReasoner, a framework that introduces novel view synthesis into the VLM reasoning process by leveraging 3D Gaussian Splatting (3DGS). Given a user query about a 3D scene, SplatReasoner retrieves relevant observations and synthesizes query-conditioned viewpoints that reveal the visual evidence needed to answer the query and ground the referred entities in 3D. Experiments show that query-conditioned novel view synthesis improves both embodied reasoning and 3D grounding over fixed-view memory and language-embedded 3DGS baselines.

URL PDF HTML 收藏
2607.09655 2026-07-13 cs.CV 新提交

OpenLongTail: Generative Scaling of Long-Tail Driving Data

OpenLongTail:长尾驾驶数据的生成式扩展

Lulin Liu, Nuo Chen, Yan Wang, Bangya Liu, Wenyan Cong, Hezhen Hu, Boris Ivanovic, Hao Wang, Ziyao Zeng, Xinyu Gong, Yang Zhou, Zixiang Xiong, Dilin Wang, Zhangyang Wang, Weisong Shi, Ruohan Zhang, Marco Pavone, Zhiwen Fan

机构 * Texas A&M University(德克萨斯农工大学) NVIDIA(英伟达) UW–Madison(威斯康星大学麦迪逊分校) UT Austin(德克萨斯大学奥斯汀分校) Yale University(耶鲁大学) Adobe(奥多比公司) Meta(元公司) University of Delaware(特拉华大学) Stanford University(斯坦福大学)

AI总结 研究针对长尾驾驶数据稀缺影响策略扩展的问题,提出开源生成数据引擎OpenLongTail,通过姿态外推视图合成管道及普吕克射线几何增强,合成异构数据提升闭环驾驶稳健性,验证了其多方面有效性。

Comments Project page: https://openlongtail.github.io/

详情
AI中文摘要

扩展稳健的驾驶策略从根本上受到策划数据集中边缘情况稀缺的限制。现实世界不断捕捉这些关键事件,但从异构源收集时,此类长尾事件仍未得到充分利用。具体而言,多样但有价值的野外长尾视频缺乏训练策略模型所需的全视图覆盖,常缺少多视图姿态或仅来自单目行车记录仪。这种模态差距阻碍了这些普遍观察结果转化为用于长尾泛化的可扩展训练数据。我们引入了OpenLongTail,一个用于在长尾事件下扩展自动驾驶策略的开源生成数据引擎。为了将异构数据源转换为对策略学习有用的视图对齐且时间连贯的多视图资产,我们开发了一个基于姿态的外推视图合成管道来生成缺失视图。我们还通过将普吕克射线几何注入可扩展生成引擎,进一步增强新生成视图的跨视图一致性和时间对齐。通过合成异构长尾数据,我们观察到在处理长尾事件时闭环驾驶稳健性有显著提高。通过测量外推视图合成和姿态指标,我们验证了OpenLongTail在视觉保真度、跨视图一致性和自我轨迹恢复方面的有效性。

英文摘要

Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.

URL PDF HTML 收藏