arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

至 收录 533
2607.18100 2026-07-21 cs.AI 新提交

Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering

我们能否使大语言模型摆脱自我循环?通过激活引导实现细粒度推理控制

Sheldon Yu, Tong Yu, Xunyi Jiang, Rohan Surana, Gagan Mundada, Sungchul Kim, Lina Yao, Julian McAuley, Junda Wu

机构 * UC San Diego(加州大学圣地亚哥分校) Adobe Research(Adobe研究院) University of New South Wales(新南威尔士大学)

AI总结 研究大语言模型推理过程不可控问题,提出SOPHIA方法,通过将推理轨迹视为潜在状态序列,构建引导向量库,在推理时进行干预,可检测并防止自我循环,提升推理质量。

详情
AI中文摘要

扩展推理已成为前沿大语言模型(LLMs)的标准,但模型生成的轨迹在很大程度上仍无法控制。现有塑造模型推理方式的方法是基于提示的,在输入层面操作,无法对推理过程本身进行细粒度控制。相关工作分析并发现了大语言模型推理轨迹中的潜在转换动态。在此基础上,我们对这些状态进行统计表征,发现失败轨迹会陷入自我循环。为干预这些失败,我们提出了SOPHIA:通过隐藏状态干预和激活来引导推理过程。我们将每个推理轨迹视为潜在状态序列,分类前缀到潜在状态,记录步骤级转换,构建引导向量库。在推理时,控制器推断当前状态,给定目标状态检索相应向量,还能从转换结构中在线检测自我循环。实验表明,我们的方法能可靠干预自我循环失败,引导向量可推广到不同状态对,细粒度可控性带来更好的推理质量。

英文摘要

Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable. Existing methods for shaping how a model reasons are prompt based approaches and operate at the input level, offering no fine-grained control over the reasoning process itself. Related work analyzes and discovers latent transition dynamics in the reasoning traces from Large Language Models. Building on this, we statistically characterize these states, and show that failure trajectories get stuck in self-loops, exhausting the token budget without progress toward the final answer. To intervene on these failures, We propose SOPHIA: Steering Of reasoning Processes via Hidden-state Intervention and Activations. We treat each reasoning trace as a sequence of latent states rather than an unstructured texts, and investigate whether inference time interventions can provide fine-grained control over the self-looping reasoning process. We classify every prefix to a latent state, record step level transitions, and use them to construct a bank of steering vectors indexed by state pairs. At inference time, a controller infers the current state and, given a target state, retrieves the corresponding vector and can also detect self-loops online from the transition structure to prevent the model from sinking into a reasoning black hole. Through extensive experiments, our method reliably intervenes on self-loop failures, with steering vectors that generalize to different state pairs. End task accuracy and token efficiency indicate that fine-grained controllability results in better reasoning quality.

URL PDF HTML 收藏
2607.17411 2026-07-21 cs.GR cs.LG 新提交

Feature-Guided Diffusion for Non-Differentiable Inverse Rendering

用于不可微逆渲染的特征引导扩散

Andrei-Timotei Ardelean, Michael Fischer, Tim Weyrich, Tomáš Iser

机构 * Adobe Research(Adobe研究)

AI总结 研究逆渲染问题,提出完全黑箱框架FIDE,通过特征引导,利用视觉Transformer提取特征训练扩散模型,结合CMA进化策略优化,在多种逆问题上验证,显著提升收敛速度并逃离局部最小值。

详情
AI中文摘要

传统上,逆渲染通过可微渲染器和梯度下降来解决,这需要大量特定问题的工程设计,并且由于模糊性容易陷入局部最小值。无导数方法减轻了工程需求,但通常严重依赖于良好的问题初始化。在这项工作中,我们提出了特征信息扩散进化(FIDE),这是一个完全黑箱的框架,不需要梯度或特定初始化:渲染器被视为一个不透明的函数,唯一的要求是生成图像。我们的关键见解是特征引导:我们不是将每个候选渲染减少到一个标量损失值,而是使用视觉Transformer(ViT)从中提取密集的视觉特征。随后,我们使用这些特征来训练基于扩散的候选提案模型,使网络能够使用视觉线索来预测与目标图像匹配的参数。然后,通过CMA进化策略在闭环中对该扩散模型提出的候选解决方案进行优化,随着优化的进行不断缩小提案区域。我们在路径追踪、向量样条、Voronoi着色器和机器人等各种逆问题上进行了验证,并证明特征引导显著提高了收敛速度,超过了标量损失基线,并可靠地逃离了基于梯度的方法停滞的局部最小值。

英文摘要

Inverse rendering is traditionally solved via differentiable renderers and gradient descent, which requires substantial problem-specific engineering and is prone to getting stuck in local minima due to ambiguities. Derivative-free approaches alleviate engineering requirements, but often heavily depend on a good problem initialization. In this work, we propose Feature-Informed Diffusion Evolution (FIDE), a fully black-box framework that requires no gradients or specific initialization: the renderer is treated as an opaque function whose only requirement is to produce images. Our key insight is feature guiding: rather than reducing each candidate rendering to a scalar loss value, we use a Vision Transformer (ViT) to extract dense visual features from it. We subsequently use these features to train a diffusion-based candidate proposal model, allowing the network to use visual cues to predict parameters that would match the target image. The candidate solutions proposed by this diffusion model are then refined in a closed loop with a CMA evolution strategy, continuously narrowing the proposal region as optimization progresses. We validate across diverse inverse problems from path tracing, vector splines, Voronoi shaders, and robotics, and demonstrate that feature-guiding substantially improves convergence over scalar-loss baselines and reliably escapes local minima where gradient-based methods stall.

URL PDF HTML 收藏
2607.16239 2026-07-21 cs.LG stat.AP 新提交

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges

BACON:用于多人工智能评判器建模与评估的预算人类校准

Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, Avi Feller, Saayan Mitra, Ritwik Sinha

机构 * Adobe Research(Adobe研究院) University of California, Berkeley(加州大学伯克利分校)

AI总结 研究针对人工智能评判器输出有偏差等问题,提出BACON四阶段流程,结合预算人类校准与多人工智能评判器输出。通过构建辅助特征、收集人类标签训练模型,实现总体指标估计和个体级替代评分,提高了预测准确性等,提供实用评估框架。

详情
AI中文摘要

人工智能评判器为人工评估提供了一种可扩展、低成本的替代方案,但其输出可能存在相对于人类偏好的偏差,且高度依赖项目,在评判器、任务和领域之间存在差异。当未校准的人工智能评估用于模型排名、项目评分或总体质量报告时,这些偏差会直接扭曲下游决策。我们提出了BACON,这是一个四阶段的流程,将预算人类校准与多个人工智能评判器的输出相结合,以产生更准确的注释。BACON为每个项目构建全覆盖的辅助特征,包括多评判器分数、令牌级不确定性统计和上下文嵌入。然后,它为一个小的采样子集收集人类标签,并训练一个交叉拟合的结果模型,以生成校准后的项目级替代预测。这些预测支持两个用例:使用具有有效置信区间的增强估计方程估计器对总体指标(如均值或分位数)进行总体估计;以及用于项目排名和注释的个体级替代评分。BACON将人工智能评判器视为辅助测量而非地面真值:人类标签提供校准锚点,而人工智能衍生的信号提高效率。在不同的任务、领域和标注预算中,BACON提高了预测准确性和排名一致性,并相对于原始人工智能输出和基于纯人类标签的方法减少了偏差和方差。这些结果表明,BACON为有限人工注释的可扩展评估提供了一个实用的、基于统计的框架。

英文摘要

AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains. When uncalibrated AI evaluations are used for model ranking, item scoring, or population-level quality reporting, these biases can directly distort downstream decisions. We propose BACON, a four-stage pipeline that combines budgeted human calibration with multiple AI-judge outputs to produce more accurate annotations. BACON constructs full-coverage auxiliary features for every item, including multi-judge scores, token-level uncertainty statistics, and contextual embeddings. It then collects human labels for a small sampled subset and trains a cross-fitted outcome model to generate calibrated item-level surrogate predictions. These predictions support two use cases: population-level estimation of summary metrics, such as means or quantiles, using an augmented estimating-equation estimator with valid confidence intervals; and individual-level surrogate scoring for item ranking and annotation. BACON treats AI judges as auxiliary measurements rather than ground truth: human labels provide the calibration anchor, while AI-derived signals improve efficiency. Across diverse tasks, domains, and labeling budgets, BACON improves predictive accuracy and ranking consistency, and reduces bias and variance relative to raw AI outputs and purely human-label-based methods. These results show that BACON offers a practical, statistically grounded framework for scalable evaluation with limited human annotation.

URL PDF HTML 收藏
2607.14645 2026-07-21 cs.CV 版本更新

Autoregressive Modeling of Film with Applications in Video Montage

电影的自回归建模及其在视频蒙太奇中的应用

Marcelo Sandoval-Castañeda, Fabian Caba Heilbron, Shiry Ginosar, Bryan Russell, Josef Sivic, Alexei A. Efros, Greg Shakhnarovich

机构 * TTI-Chicago(芝加哥丰田理工学院) Adobe(奥多比公司) Czech Institute of Informatics, Robotics and Cybernetics, Czech Technical University(捷克技术大学捷克信息学、机器人学与控制论研究所) UC Berkeley(加州大学伯克利分校)

AI总结 研究针对视频蒙太奇挑战,提出FilmGPT自回归Transformer,通过在电影语料库训练捕捉电影“语法”,推理时用镜头约束解码算法选最佳镜头,在镜头预测和电影编辑任务中表现出色,还适用于多种视频蒙太奇应用。

详情
AI中文摘要

本文介绍了FilmGPT,一种自回归Transformer,旨在应对视频蒙太奇挑战,即将原始的、“不可观看”的镜头集合转变为连贯的电影序列。受现代语言模型语言学习启发,在大量电影语料库上训练长上下文自回归Transformer,直接从数据而非手工编码规则中隐式捕捉电影“语法”。推理时引入镜头约束解码算法,根据从电影中学到的统计模式从输入原始镜头中选择最佳下一个镜头。在镜头序列排序标准基准上用FilmGPT自回归模型进行下一个镜头预测,优于先前技术水平。通过用户研究在完整电影编辑任务上评估镜头约束解码算法,基于FilmGPT的编辑显著优于先前方法。最后展示了FilmGPT在视频蒙太奇广泛应用中的适用性。

英文摘要

This work introduces FilmGPT, an autoregressive transformer designed to address the challenge of video montage--turning a collection of raw, "unwatchable" footage into coherent cinematic sequences. Inspired by language learning in modern LLMs, we train a long-context autoregressive transformer on a large corpus of movies. The aim is to implicitly capture the "grammar" of film directly from data rather than from hand-coded rules. Unlike other generative models, FilmGPT does not generate any new video frames. Instead, at inference time, we introduce a footage-constrained decoding algorithm to select the best next shot from the input raw footage according to the statistical patterns learned from films. We first evaluate these learned statistics directly by using the FilmGPT autoregressive model for next shot prediction on a standard benchmark of shot sequence ordering, outperforming the previous state of the art. We then evaluate our footage-constrained decoding algorithm on the full film editing task via a user study, and find that our FilmGPT-based editing significantly outperforms previous approaches. Finally, we demonstrate the applicability of FilmGPT to a wide range of applications in video montage, from automatic video segment trimming to human-in-the-loop film editing.

URL PDF HTML 收藏
2607.15845 2026-07-20 cs.AI 新提交

Knowledge-Centric Agents for Workflow Generation

用于工作流生成的以知识为中心的智能体

Zhendong Li, Lei Sun, Ruibo Ming, He Zhang, Danda Pani Paudel, Luc Van Gool, Jinjin Gu

机构 * INSAIT(未知机构) Sofia University “St. Kliment Ohridski”(索非亚大学“圣克莱门特·奥赫里德斯基”分校) Adobe Research(Adobe研究院)

AI总结 研究视觉创作系统中工作流生成问题,提出以知识为中心的框架,通过知识反转、注入和可逆推理进行工作流生成,实验证明该方法生成的工作流在多样性、结构连贯性和执行成功率上优于现有系统。

Comments Accepted to ECCV 2026

详情
AI中文摘要

在诸如ComfyUI等视觉创作系统中,工作流生成不仅需要句法准确性,还需要对模块化组合进行专家级推理。现有的大语言模型方法往往将其视为直接的文本到JSON生成任务,存在结构脆弱性问题,且缺乏有效设计所需的经验知识。我们认为成功的工作流生成需要对知识本身进行建模,包括其结构、层次和推理动态。为此,我们提出了一个以知识为中心的框架,该框架学习在多个抽象层次上对知识进行反转、注入和推理。我们首先进行知识反转,从大量实际工作流中提取层次表示,然后通过监督微调进行知识注入,在推理过程中,模型进行可逆推理以合成可执行工作流,并通过自我优化增强结构连贯性。大量实验表明,我们的方法比现有系统生成的工作流具有更丰富的节点多样性、更连贯的结构和更高的执行成功率,为知识驱动的智能工作流生成奠定了新基础。

英文摘要

Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large language model (LLM) approaches often treat this as a direct text-to-JSON generation task, struggling with structural brittleness and lacking the experiential knowledge required for effective design. We argue that successful workflow generation requires modeling knowledge itself, including its structure, hierarchy, and reasoning dynamics. To this end, we propose a knowledge-centric framework that learns to invert, inject, and infer with knowledge across multiple abstraction levels. We first perform knowledge inversion to distill hierarchical representations, ranging from full pseudo-codes and skeletons to high-level strategies, from large collections of real-world workflows. We then conduct knowledge injection through supervised fine-tuning, teaching the model to reason from task descriptions to strategies and from strategies to executable structures. During inference, the model performs reversible reasoning to synthesize executable workflows, augmented by self-refinement for structural coherence. Extensive experiments demonstrate that our method produces workflows with richer node diversity, more coherent structures, and higher execution success rates than existing systems, establishing a new foundation for knowledge-driven, agentic workflow generation.

URL PDF HTML 收藏
2607.15107 2026-07-17 cs.LG math.CT 新提交

Learning in Infinitesimal Non-Compositional Sketches

在无穷小非组合草图中学习

Sridhar Mahadevan

机构 * Adobe Research(Adobe研究院) University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校)

AI总结 研究如何通过LINCS框架修复机器学习中非组合性问题,将问题指定为草图,利用切提升和INC自函子,把机器学习表述为寻找余代数不动点,证明特定条件下最终INC余代数存在,正进行多场景实验评估。

详情
AI中文摘要

本文开发了一个范畴框架——无穷小非组合草图学习(LINCS),用于修复非组合性问题,即图表无法通过提升到切范畴设置的商草图进行分解。机器学习问题被指定为草图,通过通用分解问题的失败来定义非组合性。给定学习草图和模型,定义了基础缺陷和切提升,LINCS被定义为切提升后的分解障碍。本文还介绍了切学习草图,定义了INC自函子,将机器学习表述为寻找相继切展开稳定的余代数不动点。利用Aczel - Mendler定理证明了在特定条件下最终INC余代数的存在性。目前正在多个具体机器学习设置中对LINCS进行详细实验评估。

英文摘要

This paper develops a categorical framework -- Learning in Infinitesimal Non-Compositional Sketches (LINCS) -- as the repair of non-compositionality: failures of diagrams to factor through quotient sketches lifted to the tangent category setting. Machine learning problems are specified as sketches: graphs with commutativity conditions $\mathcal D$, limit cones $\mathcal L$, and colimit cocones $\mathcal K$, generalizing the usual scalarization of loss functions or vector space assumptions. Non-compositionality is defined purely as failure of a universal factorization problem, not as arithmetic error between the desired and actual predictions. Given a learning sketch $\mathbb S=(S,\mathcal D,\mathcal L,\mathcal K)$, whose underlying graph is $S$, and a model $D:J \rightarrow C$, the base defect is the obstruction to factorization $\mbox{Obs}(\mbox{Fact}_{\mathbb S}(D))$. The tangent lift applies the tangent functor $T$ to obtain $TD:J \rightarrow C$, and LINCS is defined as the obstruction $\mbox{Obs}(\mbox{Fact}_{\mathbb S}(TD))$ -- asking whether infinitesimal perturbations preserve the compositionality constraints.The paper also introduces Tangent Learning Sketches, which are sketches equipped with Cockett-Cruttwell tangent structure. The paper defines the INC endofunctor, which iterates the tangent lift, producing a tower $D,TD,T^2D, \cdots$ of factorization problems. ML is thereby formulated as the search for a coalgebraic fixed point where successive tangent unfoldings stabilize ($νT_{\mbox{INC}}$). Using the Aczel--Mendler theorem, we prove existence of a final INC coalgebra whenever $T_{\mbox{INC}}$ admits a set-based class realization that creates its final carrier. A detailed experimental evaluation of LINCS is underway in a number of concrete ML settings, including deep learning, large language models, and reinforcement learning, and is described in companion papers.

URL PDF HTML 收藏
2512.14008 2026-07-17 cs.CV 版本更新

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

稀疏-LaViDa:稀疏多模态离散扩散语言模型

Shufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin, Zijun Wei, Aditya Grover, Jason Kuen

机构 * Adobe(Adobe公司) UCLA(加州大学洛杉矶分校)

AI总结 针对掩码离散扩散模型推理速度慢的问题,提出Sparse-LaViDa框架,通过动态截断冗余掩码令牌加速采样,引入专用寄存器令牌和注意力掩码保证质量与一致性,基于LaViDa - O构建,在多任务中实现加速且保持质量。

Comments 18 pages (12 pages for the main paper and 6 pages for the appendix), 9 figures

详情
AI中文摘要

掩码离散扩散模型(MDMs)在包括图像理解、生成和编辑等广泛的多模态任务中取得了强大性能。然而,由于在每个采样步骤都需要重复处理冗余掩码令牌,其推理速度仍不理想。本文提出Sparse-LaViDa,一种新颖的建模框架,在每个推理步骤动态截断不必要的掩码令牌以加速MDM采样。为保留生成质量,引入专用寄存器令牌作为截断令牌的紧凑表示。此外,为确保训练和推理的一致性,设计了专门的注意力掩码。基于最先进的统一MDM LaViDa - O构建,Sparse-LaViDa在包括文本到图像生成、图像编辑和数学推理等各种任务中实现了高达2倍的加速,同时保持生成质量。

英文摘要

Masked Discrete Diffusion Models (MDMs) have achieved strong performance across a wide range of multimodal tasks, including image understanding, generation, and editing. However, their inference speed remains suboptimal due to the need to repeatedly process redundant masked tokens at every sampling step. In this work, we propose Sparse-LaViDa, a novel modeling framework that dynamically truncates unnecessary masked tokens at each inference step to accelerate MDM sampling. To preserve generation quality, we introduce specialized register tokens that serve as compact representations for the truncated tokens. Furthermore, to ensure consistency between training and inference, we design a specialized attention mask that faithfully matches the truncated sampling procedure during training. Built upon the state-of-the-art unified MDM LaViDa-O, Sparse-LaViDa achieves up to a 2x speedup across diverse tasks including text-to-image generation, image editing, and mathematical reasoning, while maintaining generation quality.

URL PDF HTML 收藏
2509.19244 2026-07-17 cs.CV 版本更新

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

Lavida-O:用于统一多模态理解与生成的弹性大掩码扩散模型

Shufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin, Zijun Wei, Aditya Grover, Jason Kuen

机构 * Adobe(Adobe公司) UCLA(加州大学洛杉矶分校)

AI总结 研究提出Lavida-O统一多模态MDM,采用Elastic-MoT架构,结合轻量级生成与理解分支,通过多种技术支持高效生成。该模型在多模态任务基准测试中性能领先,优于现有模型,还能加速推理,成为可扩展多模态推理和生成新范式。

Comments 31 pages, 15 figures

详情
AI中文摘要

我们提出了Lavida-O,一种用于多模态理解与生成的统一掩码扩散模型(MDM)。与现有的多模态MDM(如MMaDa和Muddit,仅支持简单图像级理解任务和低分辨率图像生成)不同,Lavida-O提供了一个单一框架,可实现图像级理解、对象定位、图像编辑和高分辨率(1024px)文本到图像合成。Lavida-O采用了新颖的弹性变压器混合(Elastic-MoT)架构,将轻量级生成分支与更大的理解分支相结合,并通过令牌压缩、通用文本条件和分层采样来支持高效和高质量的生成。Lavida-O还在图像生成和编辑任务中纳入了规划和迭代自我反思,以其理解能力无缝提升生成质量。Lavida-O在包括RefCOCO对象定位、GenEval文本到图像生成和ImgEdit图像编辑在内的广泛基准测试中取得了领先性能,优于现有的自回归模型和连续扩散模型(如Qwen2.5-VL和FluxKontext-dev),同时在推理时提供了显著的加速。这些进展使Lavida-O成为可扩展多模态推理和生成的新范式。

英文摘要

We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks and low-resolution image generation, Lavida-O presents a single framework that enables image-level understanding, object grounding, image editing, and high-resolution (1024px) text-to-image synthesis. Lavida-O incorporates a novel Elastic Mixture-of-Transformers (Elastic-MoT) architecture that couples a lightweight generation branch with a larger understanding branch, supported by token compression, universal text conditioning and stratified sampling for efficient and high-quality generation. Lavida-O further incorporates planning and iterative self-reflection in image generation and editing tasks, seamlessly boosting generation quality with its understanding capabilities. Lavida-O achieves state-of-the-art performance on a wide range of benchmarks including RefCOCO object grounding, GenEval text-to-image generation, and ImgEdit image editing, outperforming existing autoregressive models and continuous diffusion models such as Qwen2.5-VL and FluxKontext-dev, while offering considerable speedup at inference. These advances establish Lavida-O as a new paradigm for scalable multimodal reasoning and generation.

URL PDF HTML 收藏
2505.16839 2026-07-17 cs.CV 版本更新

LaViDa: A Large Diffusion Language Model for Multimodal Understanding

LaViDa:用于多模态理解的大型扩散语言模型

Shufan Li, Konstantinos Kallidromitis, Hritik Bansal, Akash Gokul, Yusuke Kato, Kazuki Kozuka, Jason Kuen, Zhe Lin, Kai-Wei Chang, Aditya Grover

机构 * UCLA(加州大学洛杉矶分校) Panasonic AI Research(松下人工智能研究) Adobe Research(Adobe研究) Salesforce Research(Salesforce研究)

AI总结 研究针对现有视觉语言模型在快速推理和可控生成方面的不足,提出基于离散扩散模型构建 LaViDa 系列视觉语言模型,采用多种新技术,在多模态基准测试中性能出色,成为自回归视觉语言模型的有力替代。

Comments 26 pages, 8 figures

详情
AI中文摘要

现代视觉语言模型(VLM)可解决各种需要视觉推理的任务。在现实场景中,VLM 理想特性包括快速推理和可控生成。现有自回归 VLM 在这些方面存在困难。离散扩散模型(DM)提供了有前景的替代方案,在多模态任务中的潜力未被充分探索。我们引入基于 DM 的 LaViDa 系列 VLM,通过为 DM 配备视觉编码器并联合微调以遵循多模态指令。LaViDa 采用了如互补掩码、前缀 KV 缓存和时间步长移位等新技术。实验表明,LaViDa 在多模态基准测试中性能优于或与自回归 VLM 竞争,具有速度 - 质量权衡灵活、可控性和双向推理等优势。

英文摘要

Modern Vision-Language Models (VLMs) can solve a wide range of tasks requiring visual reasoning. In real-world scenarios, desirable properties for VLMs include fast inference and controllable generation (e.g., constraining outputs to adhere to a desired format). However, existing autoregressive (AR) VLMs like LLaVA struggle in these aspects. Discrete diffusion models (DMs) offer a promising alternative, enabling parallel decoding for faster inference and bidirectional context for controllable generation through text-infilling. While effective in language-only settings, DMs' potential for multimodal tasks is underexplored. We introduce LaViDa, a family of VLMs built on DMs. We build LaViDa by equipping DMs with a vision encoder and jointly fine-tune the combined parts for multimodal instruction following. To address challenges encountered, LaViDa incorporates novel techniques such as complementary masking for effective training, prefix KV cache for efficient inference, and timestep shifting for high-quality sampling. Experiments show that LaViDa achieves competitive or superior performance to AR VLMs on multi-modal benchmarks such as MMMU, while offering unique advantages of DMs, including flexible speed-quality tradeoff, controllability, and bidirectional reasoning. On COCO captioning, LaViDa surpasses Open-LLaVa-Next-8B by +4.1 CIDEr with 1.92x speedup. On bidirectional tasks, it achieves +59% improvement on Constrained Poem Completion. These results demonstrate LaViDa as a strong alternative to AR VLMs. Code and models will be released in the camera-ready version.

URL PDF HTML 收藏
2607.12706 2026-07-16 cs.SD 版本更新

AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

AutoSIFT:用于可控语音生成的自动风格筛选与任意风格填充

Haowei Lou, Junda Wu, Chengkai Huang, Tong Yu, Hye-young Paik, Wen Hu, Lina Yao

机构 * UNSW Sydney(新南威尔士大学悉尼分校) University of California San Diego(加利福尼亚大学圣地亚哥分校) Macquarie University(麦考瑞大学) Adobe Research(Adobe 研究院)

AI总结 研究针对TTS模型难以细粒度控制说话风格的问题,提出AutoSIFT框架,将风格分解为可描述和残余两类,通过广义风格解缠器和任意风格填充器,可在保留残余风格时替换指定风格类别,实现高度可定制的语音生成。

详情
AI中文摘要

当前最先进的文本到语音(TTS)模型在自然度和表现力方面表现出色,但对说话风格进行细粒度、解耦控制仍具有挑战性。在电影配音、游戏语音表演和视频内容生成等专业场景中,用户常需修改特定风格类别,同时保留其他风格。现有方法难以联合控制显式语义属性并保留细微的韵律细节。我们提出AutoSIFT,一个用于类别级风格编辑的可控语音生成框架。它将说话风格分解为已知的可文本描述类别和未知的残余风格,通过广义风格解缠器和任意风格填充器,在保留残余语音风格的同时替换文本指定的风格类别,实现自然、富有表现力和高度可定制的语音生成。

英文摘要

State-of-the-art text-to-speech (TTS) models achieve impressive naturalness and expressiveness, yet fine-grained, disentangled control over speaking styles remains challenging. In professional scenarios such as film dubbing, game voice acting, and video content generation, users often need to modify a specific style category, such as emotion, age, or gender, while preserving all others. Existing style-controllable TTS methods typically rely on either text-described styles or speech-reference style transfer, making it difficult to jointly control explicit semantic attributes and preserve subtle, text-undescribed prosodic details. We propose AutoSIFT, a controllable speech generation framework for category-level style editing. AutoSIFT decomposes speaking style into known text-describable categories and unknown residual styles that capture non-verbal prosody and speaker-specific nuances. It consists of a generalized Style Disentangler, which extracts category-aware style prototypes from reference speech, and an Arbitrary Style Infiller, which selectively infills unspecified style categories from the reference. By replacing only text-specified style categories while preserving residual speech-derived styles, AutoSIFT enables natural, expressive, and highly customizable speech generation.

URL PDF HTML 收藏
2607.11205 2026-07-14 cs.CV 新提交

Parallax Portrait Matting

视差人像抠图

Xin Cai, Jiawen Chen, Lars Jebe, Tianfan Xue, Zhoutong Zhang

机构 * Multimedia Laboratory, The Chinese University of Hong Kong(香港中文大学多媒体实验室) Adobe NextCam Shanghai AI Laboratory(上海人工智能实验室) CPII under InnoHK(创新香港研发平台下的CPII)

AI总结 针对图像抠图难题,提出视差人像抠图方法,利用连拍摄影中相机小幅度运动产生的前景-背景视差,通过估计trimap和前景/背景运动构建对齐视图预测,能恢复更精细细节和准确前景颜色。

Comments ECCV 2026

详情
AI中文摘要

图像抠图是病态问题,在前景和背景纹理丰富时尤其困难。单图像抠图方法虽能从数据中学到强先验,但在复杂情况下表现不佳。现有方法需额外信号如绿幕、偏振光或干净背景图像来改善结果,且通常依赖特殊拍摄设置。我们提出视差人像抠图,一种实用的双帧抠图方法,利用轻微视角变化拍摄的第二张图像。此设置在连拍摄影中自然出现,相机小幅度运动产生前景-背景视差并为抠图提供补充观测。我们的流程估计trimap和前景/背景运动,构建对齐视图用于预测。为处理不完美的运动估计,网络使用背景对齐对直接融合,通过交叉注意力利用前景对齐线索进行误差补偿。实验表明,在具有挑战性的人像案例中,我们的方法比强大的单图像抠图基线能恢复更精细的细节和更准确的前景颜色。

英文摘要

Image matting is highly ill-posed, especially when both the foreground and background are richly textured. While single-image matting methods learn strong priors from data, they often struggle on these challenging cases. Existing approaches improve results by requiring additional signals such as green screens, polarized lighting, or clean background images, but these typically rely on specialized capture setups. We present Parallax Portrait Matting, a practical two-frame matting method that uses a second image captured with slight viewpoint change. Such a setting arises naturally in burst photography, where small camera motion induces foreground-background parallax and provides complementary observations for matting. Our pipeline estimates trimaps and foreground/background motion, then constructs aligned views for prediction. To handle imperfect motion estimation, the network uses the background-aligned pair for direct fusion and the foreground-aligned cue through cross-attention for error compensation. Experiments show that our method recovers finer details and more accurate foreground colors than strong single-image matting baselines on challenging portrait cases.

URL PDF HTML 收藏
2607.10463 2026-07-14 cs.AI cs.IR 新提交

GRASP: GRanularity-Aware Search Policy for Agentic RAG

GRASP:用于智能检索增强生成的粒度感知搜索策略

Varun Gandhi, Jaewook Lee, Shantanu Todmal, Franck Dernoncourt, Ryan Rossi, Zichao Wang, Andrew Lan

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Adobe Research(Adobe 研究院)

AI总结 研究智能体检索增强生成中模型决策难题,提出GRASP强化学习框架,训练智能体协调互补检索工具,实验表明其提升检索召回率和问答性能,还展现出可解释行为,凸显协调检索信号和上下文粒度对智能体推理的关键作用。

详情
AI中文摘要

智能检索增强生成(Agentic RAG)通过允许语言模型迭代推理、生成搜索查询、检索证据和预测答案来扩展静态RAG。然而,模型在决定何时检索、使用词汇匹配还是语义相似性以及如何控制上下文粒度以防止无关令牌干扰智能体推理方面仍然具有挑战性。本文介绍了GRASP,这是一个强化学习框架,用于训练智能体在多步推理过程中自适应地协调互补检索工具。GRASP为智能体提供语义搜索、关键词搜索和段落阅读操作,使其能够仅在需要时检索句子级证据并扩展进一步的上下文。我们使用一种联合考虑答案准确性、基于事实的阅读、互补搜索和轮次效率的奖励来训练策略。在多跳推理基准上的实验表明,与单步检索、基于提示的智能体RAG和基于强化学习的检索基线相比,GRASP提高了检索召回率和下游问答性能。定性和消融分析表明,学习到的策略发展出了可解释的略读和扫描行为:它使用语义搜索进行广泛探索,段落阅读进行局部验证,关键词搜索用于特定实体的证据。这些结果表明,学习协调检索信号和上下文粒度对于智能体的正确推理至关重要。

英文摘要

Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when to retrieve, whether to use lexical matching or semantic similarity, and how to control context granularity to prevent irrelevant tokens from interfering with agent reasoning. In this paper, we introduce GRASP, a reinforcement learning (RL) framework for training agents to adaptively coordinate complementary retrieval tools during multi-step reasoning. GRASP provides the agent with semantic search, keyword search, and paragraph-reading actions, enabling it to retrieve sentence-level evidence and expand further context only when needed. We train the policy with a reward that jointly accounts for answer accuracy, grounded reading, complementary search, and turn efficiency. Experiments on multi-hop reasoning benchmarks show that GRASP improves both retrieval recall and downstream question answering performance compared with single-step retrieval, prompting-based agentic RAG, and RL-based retrieval baselines. Qualitative and ablation analyses show that the learned policy develops interpretable skimming and scanning behavior: it uses semantic search for broad exploration, paragraph reading for local verification, and keyword search for entity-specific evidence. These results suggest that learning to coordinate retrieval signals and context granularity is critical for agent's correct reasoning.

URL PDF HTML 收藏
2607.09794 2026-07-14 cs.AI cs.MA 新提交

Agentic Context Learning with Self-Discovered Specification

通过自我发现的规范进行智能体上下文学习

Jike Zhong, Ming Li, Yuxiang Lai, Ziyan Yang, Jingyu Xie, Jihyung Kil, Zheda Mai, Shao-Yuan Lo, Ren Xiang, Konstantinos Psounis, Yuanyuan Lei

机构 * University of Southern California(南加州大学) University of Florida(佛罗里达大学) Emory University(埃默里大学) Adobe Research(Adobe研究院) The Ohio State University(俄亥俄州立大学) National Taiwan University(国立台湾大学)

AI总结 研究上下文学习困难的原因,发现其不仅需内容获取还需规范获取。设计PSCI干预措施,提取并强化局部规范,在多个模型上取得显著提升,验证了规范获取的重要性,表明上下文学习依赖内容与规范获取。

详情
AI中文摘要

上下文学习是一种新兴的推理时任务,大型语言模型(LLMs)必须从预训练中不存在的复杂上下文中学习并应用新颖的、特定于任务的知识;即使是前沿模型的任务成功率也低于24%。本文进行了全面的实证研究以理解为何此设置仍很困难。一个自然假设是失败源于内容获取,但在CL - Bench这个广泛的上下文学习基准上的十二个检索、反思和验证基线中,我们发现相较于直接的全上下文提示,收益有限。进一步的失败分析揭示,与典型的长上下文任务不同,上下文学习不仅需要恢复局部内容,还需要获取局部规范,这些规范在查询中通常未明确指定但分布在上下文中。在所有31592个评分项目中,55.4%明确评估规范获取,而只有22.6%评估内容获取。此外,尽管76.7%的规范在用户查询中未指定,但95.5%可追溯到上下文,表明这些是可学习的义务而非隐藏要求。为验证此诊断,我们设计了一个故意简单的干预措施PSCI(私有规范 - 合同归纳),它提取局部规范并通过对抗性检查和修复来执行;在CL - Bench上,PSCI使用GPT - 5.1达到了28.14%的最新水平(绝对提高5.59个百分点,相对提高24.8%),在Qwen3.5 - 27B(提高5.28个百分点)和Gemini 3 Pro(提高6.17个百分点)上也得到了复制。十七个消融实验进一步分离了特定于任务的规范的作用。总体而言,我们的结果表明上下文学习不仅取决于内容获取,还取决于规范获取。

英文摘要

Context learning is an emerging inference-time task where LLMs must learn and apply novel, task-specific knowledge from intricate contexts absent from pre-training; even frontier models score under 24% task success. In this work, we conduct a comprehensive empirical study to understand why this setting remains difficult. A natural hypothesis is that failures stem from content access; yet across twelve retrieval, reflection, and verification baselines on CL-Bench, an extensive context learning benchmark, we find limited gains over direct full-context prompting. Further failure analysis reveals a key finding: unlike typical long-context tasks such as long document understanding, context learning requires not only recovering local content but also acquiring local specifications that are often unspecified in the query but distributed across the context: domain-specific formats, local rules, and completeness conditions. Across all 31,592 rubric items, we find that 55.4% clearly evaluate specification acquisition, while only 22.6% evaluate content acquisition. Moreover, despite 76.7% of specifications being unspecified in the user query, 95.5% are traceable to the context, indicating these are learnable obligations rather than hidden requirements. To validate this diagnosis, we design a deliberately simple intervention PSCI (private specification-contract induction) which extracts local specifications and enforces them through adversarial checking and repair; PSCI achieves state-of-the-art 28.14% with GPT-5.1 (+5.59 pp absolute and +24.8% relative) on CL-Bench, replicated on Qwen3.5-27B (+5.28 pp) and Gemini 3 Pro (+6.17 pp). Seventeen ablations further isolate the role of task-specific specifications. Overall, our results suggest context learning hinges on not only content acquisition but also specification acquisition.

URL PDF HTML 收藏
2607.09665 2026-07-14 cs.AI cs.CL cs.LG 新提交

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

格式敏感性指数:大语言模型基准测试中令牌控制的提示包装器稳健性与模式合规性

Deep Pankajbhai Mehta

机构 * Adobe Inc(Adobe公司)

AI总结 研究提示包装器格式差异对模型分数的影响,引入格式敏感性指数(FSI)和可解析性敏感性指数(PSI),通过大量实验发现平均FSI在不同模型间变化超30倍,且可解析性是准确性有力预测指标,指出不考虑包装器差异和合规性报告准确性不可靠并给出建议。

Comments 10 pages, 6 figures

详情
AI中文摘要

提示包装器通常仅在格式上有所不同,但它们可以使模型分数变化到足以改变排行榜结论。我们在令牌控制协议下研究这种差异,并引入两个互补指标:格式敏感性指数(FSI),即由包装器选择引起的准确性范围;以及可解析性敏感性指数(PSI),即答案可解析性的相应范围。在跨越7个问答任务、5个包装器家族以及4个参数从7B到72B的指令模型的140,000次OpenRouter生成中,我们发现平均FSI在不同模型间变化超过30倍,且很大程度上由合规失败导致。固定效应回归表明,即使在控制任务、模型和包装器后,可解析性仍是准确性的有力预测指标。我们认为,在不报告包装器差异和合规性的情况下报告准确性在统计上是不可靠的,并为基准测试和结构化输出部署给出了实际建议。

英文摘要

Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions. We study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced by wrapper choice, and the Parseability Sensitivity Index (PSI), the corresponding range in answer parseability. Across 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by over 30x across models and is largely explained by compliance failures. A fixed-effects regression shows that parseability remains a strong predictor of accuracy even after controlling for task, model, and wrapper. We argue that reporting accuracy without wrapper variance and compliance is statistically fragile, and we give practical recommendations for both benchmarking and structured-output deployments.

URL PDF HTML 收藏
2607.11493 2026-07-14 cs.LG cs.AI math.CT 新提交

Agentic Skill Optimization over Lie Algebroids

李代数胚上的智能体技能优化

Sridhar Mahadevan

机构 * Adobe Research(Adobe研究院) University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校)

AI总结 研究智能体技能优化问题,提出LASKO框架,将技能建模为李代数胚截面,利用李括号筛选测试替代昂贵验证,在自然语言任务因果提取中实现近15倍加速,提升技能优化速度。

Comments 20 pages

详情
AI中文摘要

智能体系统越来越多地通过编辑技能来自我提升,如提示、评分标准、计划等。技能编辑并非向量空间中的独立坐标,其效果需在部署、验证和评估后才能观察到,不同编辑可能有相同即时可见效果但在其他方面存在差异,编辑顺序也很重要。本文介绍了一种新的技能优化框架LASKO(李代数胚技能优化)。它将带类型、锚定的Markdown技能建模为基础范畴,可用编辑策略作为带锚ρ的受控李代数胚的截面。锚将编辑策略映射到其可见的Markdown效果,核ker(ρ)表示潜在模板等结构,代数胚括号衡量非交换编辑组合。初步基准结果表明,LASKO在技能优化中实现了数量级的加速,主要是因为它在进行昂贵的验证前,先用微秒级的李括号筛选测试。在自然语言任务的因果提取中,与通过运行671B参数的DeepSeek V3.1 4位模型验证所有编辑的暴力方法相比,LASKO实现了近15倍的加速。

英文摘要

Agentic systems increasingly improve themselves by editing skills: prompts, rubrics, plans, tool contracts, examples, validators, and traces. Skill edits are not independent coordinates in a vector space: they are local repairs to structured artifacts whose effects are observed only after rollout, validation, and critique. Distinct edits can have the same immediate visible effect while differing in routing context, template state, guardrail scope, or future composability. The order of edits can matter as well: repairing a schema before a normalization rule need not be equivalent to applying the same edits in the reverse order. This paper introduces a new framework for skill optimization called LASKO, for Lie Algebroid SKill Optimization. LASKO models typed, anchored Markdown skills as the base category and available edit policies as sections of a controlled Lie algebroid with anchor $ρ$. The anchor maps an edit policy to its visible Markdown effect; the kernel $\ker(ρ)$ represents latent template, routing, or implementation structure; and the algebroid bracket measures noncommuting edit composition. As shown in the paper, LASKO achieves order-of-magnitude speedups in skill optimization in our preliminary benchmark results, primarily because it substitutes inexpensive Lie-bracket screening tests that run in microseconds, before investing in expensive validations that require running large language models. On a causal extraction from natural language task, LASKO achieved a speedup of almost $15 \times$ compared to a brute-force approach that validated all edits by running them through a DeepSeek V3.1 4-bit model with 671B parameters.

URL PDF HTML 收藏
2407.06150 2026-07-14 cs.CV 版本更新

PanDORA: Casual HDR Radiance Acquisition of Indoor Scenes for Image-based Lighting

PanDORA: 用于基于图像的照明的室内场景因果高动态范围辐射获取

Mohammad Reza Karimi Dastjerdi, Dominique Tanguay-Gaudreau, Frédéric Fortier-Chouinard, Yannick Hold-Geoffroy, Nima Kalantari, Jean-François Lalonde

机构 * Université Laval(洛瓦尔大学) Adobe(Adobe公司) Texas A&M University(德克萨斯农工大学)

AI总结 PanDORA通过双摄像头和NeRF算法高效获取高动态范围辐射数据,提升室内场景真实光照重建效果。

Comments Accepted at ICCP 2026. Selected for PAMI Special Issue

详情
AI中文摘要

大多数新型视图合成方法,包括神经辐射场(NeRF),难以捕捉高动态范围(HDR)辐射以实现逼真的基于图像的照明(IBL)。这种限制源于对低动态范围(LDR)图像的依赖,无法捕捉室内环境中发现的光源强度。虽然曝光序列可以恢复此范围,但通常对于实际大规模采集来说太慢。在本文中,我们介绍了PanDORA:PANoramic Dual-Observer Radiance Acquisition,一种专门设计用于快速且经济地获取高质量HDR辐射图的系统。我们的方法利用两个安装在便携单脚架上的360度摄像头同时记录不同曝光的视频。这些视频通过我们提出的两阶段NeRF基算法进行处理,该算法包含一个新颖的自校准管道来估计相机参数。此管道产生非饱和的HDR辐射场,准确捕捉场景的辐射。在新的真实室内环境HDR真实光照数据集上评估时,PanDORA在重建下游渲染任务所需的峰值强度方面表现出色,提供了一种可扩展且高效的捕捉真实世界IBL的解决方案。

英文摘要

Most novel view synthesis methods -- including Neural Radiance Fields (NeRF) -- struggle to capture the high dynamic range (HDR) radiance required for realistic image-based lighting (IBL). This limitation stems from a reliance on low dynamic range (LDR) imagery, which fails to capture the intensity of light sources found in indoor environments. While exposure bracketing can recover this range, it is often too slow for practical, large-scale acquisition. In this work, we introduce PanDORA: PANoramic Dual-Observer Radiance Acquisition, a system specifically designed for the fast and affordable capture of high-quality HDR radiance maps for IBL. Our approach utilizes two 360° cameras mounted on a portable monopod to simultaneously record videos at different exposures. These videos are processed by our proposed two-stage NeRF-based algorithm featuring a novel self-calibrating pipeline to estimate camera parameters. This pipeline produces non-saturated HDR radiance fields that accurately capture the radiance of a scene. When evaluated on a new dataset of real indoor environments featuring HDR ground truth lighting, PanDORA demonstrates superior fidelity in reconstructing the peak intensities necessary for downstream rendering tasks, providing a scalable and efficient solution for capturing real-world IBLs.

URL PDF HTML 收藏
2604.07643 2026-07-14 cs.HC cs.CL 交叉投稿

Narrix: Remixing Narrative Strategies from Examples for Story Writing

Narrix:从示例中 remix 叙事策略以进行故事写作

Chao Zhang, Shunan Guo, Abe Davis, Eunyee Koh

机构 * Cornell University(康奈尔大学) Adobe Research(Adobe研究)

AI总结 Narrix 通过可视化叙事策略帮助新手作家提高叙事能力,通过交互式故事弧和颜色编码提示,使作家能有效复用策略,提升创作信心和适应能力。

Comments 24 pages, 10 figures. To appear in CHI '26: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, April 13-17, 2026, Barcelona, Spain. DOI: https://doi.org/10.1145/3772318.3790813

详情
AI中文摘要

有经验的讲故事者将故事分解为局部叙事策略及其如何塑造更高层次的弧线。这种分解帮助作家识别他人作品中的模式并将其应用于新故事。然而,新手难以识别或有效复用这些策略。我们提出了 Narrix,一种新的写作工具,帮助新手作家识别示例故事中的叙事策略并将其应用于自己的写作。Narrix 分析示例故事中的策略,用颜色编码的词汇提示和解释突出显示,并将其置于可交互的故事弧上,通过情感变化和转折点进行探索。作家然后可以将策略拖到多维轨道上,并通过受指定策略控制的生成进行修订或继续草稿。通过一项受试者内研究(N=12),Narrix 显示相比基于聊天的写作界面,参与者在保留、信心和叙事策略适应能力方面有所提高。

英文摘要

Experienced storytellers decompose stories into local narrative strategies and how these strategies shape higher-level arcs. This decomposition helps writers recognize patterns in others' work and adapt those patterns to tell new stories. Novices, however, struggle to identify these strategies or to reuse them effectively. We present Narrix, a novel writing tool that helps novice writers recognize narrative strategies in example stories and repurpose these strategies in their own writing. Narrix analyzes strategies in example stories, highlights them with color-coded lexical cues and explanations, and situates them on an interactive story arc for exploration by emotional shifts and turning points. Writers then drag strategies onto multi-dimensional tracks and apply block-scoped edits to revise or continue their drafts through controlled generation steered by specified strategies. Through a within-subjects study (N=12), Narrix showed improved participants' retention, confidence, and creative adaptation of narrative strategies compared to a baseline chat-based writing interface.

URL PDF HTML 收藏
2506.21833 2026-07-14 cs.LG 版本更新

Memory Savings at What Cost? A Study of Alternatives to Backpropagation

以何种代价节省内存?反向传播替代方案研究

Kunjal Panchal, Sunav Choudhary, Yuriy Brun, Hui Guan

机构 * College of Information and Computer Sciences, University of Massachusetts, Amherst(信息与计算机科学学院,马萨诸塞大学阿默斯特分校) Adobe, San Jose(Adobe公司,圣乔斯)

AI总结 研究比较BP、检查点BP、FmAD和ZO在LLM及视觉语言模型训练中的表现,发现FmAD和ZO虽省内存但有计算成本等问题,带检查点的BP在多方面更优,纠正了此前关于LLM优化的片面基准测试观点。

Comments Accepted to ICML 2026

详情
AI中文摘要

前向模式自动微分(FmAD)和零阶(ZO)优化越来越多地被提议作为大语言模型(LLM)微调中节省内存且无需反向传播的替代方案。但以往通常仅将它们与标准反向传播(BP)评估,忽略了如激活检查点等节省内存的变体。本文对BP、检查点BP、FmAD和ZO在LLM及视觉语言模型训练中进行了统一理论和实证比较。结果表明,FmAD和ZO虽减少激活内存,但以更高计算成本和更长收敛时间为代价,导致准确率降低和训练变慢,尤其是在受限扰动预算下。跨模型来看,带检查点的BP优于FmAD和ZO变体,在可比内存使用下,准确率高31.1%,收敛快34.8%,计算量少3.8倍,还揭示了FmAD和ZO中与不稳定性相关的失败模式。总体而言,研究结果纠正了片面的基准测试观点,表明节省内存的方法存在根本不同的权衡,忽略这些区别会导致先前工作中关于LLM优化的误导性结论。

英文摘要

Forward-mode automatic differentiation (FmAD) and zero-order (ZO) optimization are increasingly proposed as memory-efficient, backpropagation-free alternatives for large language model (LLM) fine-tuning, yet their benefits are typically evaluated only against standard backpropagation (BP), omitting memory-efficient variants such as activation checkpointing. We present a unified theoretical and empirical comparison of BP, checkpointed BP, FmAD, and ZO for LLM and vision-language model training, showing that while FmAD and ZO reduce activation memory, they trade memory for higher computational cost and longer wall-clock time to convergence, resulting in lower accuracy and slower training, especially under constrained perturbation budgets. Across models, BP with checkpointing outperforms FmAD and ZO variants, including variance-reduced methods, achieving up to 31.1% higher accuracy, 34.8% faster convergence, and 3.8x fewer computations at comparable memory usage, while also revealing instability-related failure modes in FmAD and ZO. Overall, our results correct a one-sided benchmarking narrative by showing that memory-efficient methods entail fundamentally different trade-offs, and that ignoring these distinctions has led to misleading conclusions about LLM optimization in prior work. Our source code is available at {https://github.com/Astuary/Gradient_Estimation_Methods}.

URL PDF HTML 收藏
2607.09655 2026-07-13 cs.CV 新提交

OpenLongTail: Generative Scaling of Long-Tail Driving Data

OpenLongTail:长尾驾驶数据的生成式扩展

Lulin Liu, Nuo Chen, Yan Wang, Bangya Liu, Wenyan Cong, Hezhen Hu, Boris Ivanovic, Hao Wang, Ziyao Zeng, Xinyu Gong, Yang Zhou, Zixiang Xiong, Dilin Wang, Zhangyang Wang, Weisong Shi, Ruohan Zhang, Marco Pavone, Zhiwen Fan

机构 * Texas A&M University(德克萨斯农工大学) NVIDIA(英伟达) UW–Madison(威斯康星大学麦迪逊分校) UT Austin(德克萨斯大学奥斯汀分校) Yale University(耶鲁大学) Adobe(奥多比公司) Meta(元公司) University of Delaware(特拉华大学) Stanford University(斯坦福大学)

AI总结 研究针对长尾驾驶数据稀缺影响策略扩展的问题,提出开源生成数据引擎OpenLongTail,通过姿态外推视图合成管道及普吕克射线几何增强,合成异构数据提升闭环驾驶稳健性,验证了其多方面有效性。

Comments Project page: https://openlongtail.github.io/

详情
AI中文摘要

扩展稳健的驾驶策略从根本上受到策划数据集中边缘情况稀缺的限制。现实世界不断捕捉这些关键事件,但从异构源收集时,此类长尾事件仍未得到充分利用。具体而言,多样但有价值的野外长尾视频缺乏训练策略模型所需的全视图覆盖,常缺少多视图姿态或仅来自单目行车记录仪。这种模态差距阻碍了这些普遍观察结果转化为用于长尾泛化的可扩展训练数据。我们引入了OpenLongTail,一个用于在长尾事件下扩展自动驾驶策略的开源生成数据引擎。为了将异构数据源转换为对策略学习有用的视图对齐且时间连贯的多视图资产,我们开发了一个基于姿态的外推视图合成管道来生成缺失视图。我们还通过将普吕克射线几何注入可扩展生成引擎,进一步增强新生成视图的跨视图一致性和时间对齐。通过合成异构长尾数据,我们观察到在处理长尾事件时闭环驾驶稳健性有显著提高。通过测量外推视图合成和姿态指标,我们验证了OpenLongTail在视觉保真度、跨视图一致性和自我轨迹恢复方面的有效性。

英文摘要

Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.

URL PDF HTML 收藏
2606.00726 2026-07-13 cs.AI 版本更新

Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs

潜在奖励引导:一种自适应推理时框架,隐式促进推理大语言模型中的认知行为

Jiakang Li, Guanyu Zhu, Can Jin, Chenxi Huang, Dexu Yu, Ronghao Chen, Yang Zhou, Hongwu Peng, Xuanqi Lan, Dimitris N. Metaxas, Youhua Li

机构 * Rutgers University(罗格斯大学) South China Agricultural University(华南农业大学) Columbia University(哥伦比亚大学) Fenz.AI QuantaAlpha Adobe Santa Clara University(圣克拉拉大学) City University of Hong Kong(香港城市大学)

AI总结 提出潜在奖励引导(LRS)框架,通过优化稀疏自编码器潜在状态隐式促进认知行为,利用最终答案正确性训练潜在奖励模型估计中间状态质量,并在推理时提供状态特定的修正方向,实验表明该方法能提升推理性能并修复原始推理错误。

详情
AI中文摘要

强推理不仅依赖于模型知识,还取决于生成过程中认知行为的有效部署。现有方法通常依赖显式的行为级控制,当失败和所需修正因推理状态、任务和模型而异时,其适应性不足。为此,我们提出潜在奖励引导(LRS),一种自适应推理时框架,通过优化隐式携带认知行为的稀疏自编码器(SAE)潜在状态来促进认知行为。LRS不依赖预定义的认知行为或由此衍生的引导方向,而是基于最终答案正确性在推理轨迹上训练潜在奖励模型,以估计中间潜在状态的质量。推理时,奖励梯度为脆弱的潜在状态提供状态特定的修正方向,而奖励与置信度门控将干预限制在奖励信号标记为脆弱的状态上。在多个推理LLM骨干和基准上的实验表明,LRS一致地提升了相对于各种基线的性能,事后分析进一步表明LRS隐式促进了修复原始推理错误的良好认知行为。代码见:https://github.com/jiakanglee/Latent-Reward-Steering。

英文摘要

Strong reasoning depends not only on model knowledge but also on how effectively cognitive behaviors are deployed during generation. Existing methods often rely on explicit behavior-level control, making them insufficiently adaptive when failures and required corrections vary across reasoning states, tasks, and models. To this end, we propose Latent Reward Steering (LRS), an adaptive inference-time framework that promotes cognitive behaviors by optimizing the sparse-autoencoder (SAE) latent states that implicitly carry them. Rather than relying on predefined cognitive behaviors or steering directions derived from them, LRS trains a latent reward model on reasoning traces by final answer correctness to estimate the quality of intermediate latent states. During inference, reward gradients provide state-specific correction directions for fragile latent states, while a reward and confidence gate restricts intervention to states the reward signal flags as fragile. Experiments on multiple reasoning LLM backbones and benchmarks show that \ours consistently improves performance over various baselines, and post-hoc analyses further indicate that \ours implicitly promotes good cognitive behaviors that fix the original reasoning errors. Code is available at: https://github.com/jiakanglee/Latent-Reward-Steering.

URL PDF HTML 收藏
2607.06701 2026-07-09 cs.CV cs.AI cs.GR cs.LG cs.RO 新提交

SPEAR: A Simulator for Photorealistic Embodied AI Research

SPEAR:用于逼真的具身人工智能研究的模拟器

Mike Roberts, Renhan Wang, Rushikesh Zawar, Rachith Dey-Prakash, Quentin Leboutet, Stephan R. Richter, Matthias Müller, German Ros, Rui Tang, Stefan Leutenegger, Yannick Hold-Geoffroy, Kalyan Sunkavalli, Vladlen Koltun

机构 * Adobe Research(Adobe研究院) Intel Labs(英特尔实验室) Manycore Tech Inc(众核科技公司) Adobe(Adobe公司) NVIDIA(英伟达公司) ETH Zurich(苏黎世联邦理工学院) Imperial College London(伦敦帝国学院)

AI总结 研究针对现有逼真模拟器局限,提出SPEAR这一Python库,通过模块化插件架构连接控制UE应用程序,提升可编程性与渲染速度,还提供新图像模态和高级编程模型,经多样示例应用展示了其效用。

Comments Accepted for publication at the European Conference on Computer Vision (ECCV) 2026

详情
AI中文摘要

交互式模拟器已成为训练具身智能体和生成合成视觉数据的强大工具,但现有的逼真模拟器在通用性、可编程性和渲染速度方面存在局限。本文引入了SPEAR来解决这些问题。SPEAR是一个Python库,通过模块化插件架构连接并可编程控制任何虚幻引擎(UE)应用程序,将超14K独特UE功能暴露给Python,大幅提升可编程功能。它能以73帧每秒的速度将1920x1080逼真的美图像直接渲染到用户的NumPy数组中,比现有UE插件快一个数量级,还提供现有模拟器没有的真实图像模态。此外,SPEAR引入了一种表达性强的高级编程模型。通过各种示例应用展示了SPEAR的效用,如控制多个具身智能体、渲染城市规模环境等。

英文摘要

Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We address these limitations by introducing SPEAR: A Simulator for Photorealistic Embodied AI Research. At its core, SPEAR is a Python library that can connect to, and programmatically control, any Unreal Engine (UE) application via a modular plugin architecture. SPEAR exposes over 14K unique UE functions to Python, representing an order-of-magnitude increase in programmable functionality over existing UE-based simulators. Additionally, a single SPEAR instance can render 1920x1080 photorealistic beauty images directly into a user's NumPy array at 73 frames per second - an order of magnitude faster than existing UE plugins - while also providing ground truth image modalities that are not available in any existing UE-based simulator (e.g., a non-diffuse intrinsic image decomposition, material IDs, and physically based shading parameters). Finally, SPEAR introduces an expressive high-level programming model that enables users to specify complex graphs of UE work with arbitrary data dependencies among work items, and to execute these graphs deterministically within a single UE frame. We demonstrate the utility of SPEAR through a diverse collection of example applications: controlling multiple embodied agents with distinct action spaces (e.g., humans, cars, and robots) across several in-the-wild UE projects; rendering photorealistic city-scale environments; manipulating UE's procedural content generation systems; rendering synchronized multi-view images of detailed human faces; coordinating an interactive co-simulation with the MuJoCo physics simulator; and editing scenes with natural language via an AI coding assistant.

URL PDF HTML 收藏
2607.06527 2026-07-08 cs.CL cs.AI 新提交

RSF-GLLM: Bridging the Semantic Gap in Multi-Hop Knowledge Graph QA via Recurrent Soft-Flow and Decoupled LLM Generation

RSF-GLLM:通过循环软流和去耦语言模型生成弥合多跳知识图谱问答中的语义鸿沟

Sambaran Bandyopadhyay, Ananth Muppidi

机构 * Adobe Research(Adobe研究院) Adobe Systems(Adobe系统公司)

AI总结 研究多跳知识图谱问答中传统方法的语义鸿沟问题,提出RSF-GLLM框架,通过循环软流模块传播分数、引入正则化保证收敛,提取路径微调LLM,实验证明该方法性能优且推理效率高。

Comments Accepted for publication in ICML 2026 as a full research paper; 21 pages

详情
AI中文摘要

多跳知识图谱问答面临关键挑战:传统的先检索后读取管道破坏了可微性,使检索器难以弥合中间节点与查询缺乏词汇重叠的语义鸿沟。为解决此问题,我们提出了RSF-GLLM,一个将可微图推理与答案生成解耦的框架。我们的循环软流(RSF)模块采用GRU引导的查询更新器来传播连续相关分数,利用动态门控机制通过结构线索遍历语义不同的桥梁节点。我们引入流稀疏正则化从理论上保证从软概率到离散推理路径的收敛。这些路径被提取并文本化以微调大型语言模型(LLM),确保生成基于事实拓扑。在WebQSP和CWQ上的实验表明,与基于LLM的计算昂贵方法相比,RSF-GLLM实现了具有竞争力的性能和更高的推理效率。

英文摘要

Multi-hop Question Answering over Knowledge Graphs faces a critical challenge: traditional retrieve-then-read pipelines break differentiability, preventing the retriever from learning to bridge the semantic gap where intermediate nodes lack lexical overlap with the query. To address this, we propose RSF-GLLM, a framework decoupling differentiable graph reasoning from answer generation. Our Recurrent Soft-Flow (RSF) module employs a GRU-guided query updater to propagate continuous relevance scores, utilizing a dynamic gating mechanism to traverse semantically dissimilar bridge nodes via structural cues. We introduce flow sparsity regularization to theoretically guarantee convergence from soft probabilities to discrete reasoning paths. These paths are extracted and textualized to fine-tune a Large Language Model (LLM), ensuring generation is grounded in factual topology. Experiments on WebQSP and CWQ demonstrate that RSF-GLLM achieves competitive performance with superior inference efficiency compared to LLM based computationally expensive approaches.

URL PDF HTML 收藏
2607.05938 2026-07-08 cs.GR cs.RO 新提交

Prior-First, Condition-Second: Scalable and Controllable Hand Motion Completion

先验优先,条件其次:可扩展且可控的手部运动补全

Mingyi Shi, Xuelin Chen, Taku Komura

机构 * The University of Hong Kong(香港大学) Adobe Research(Adobe研究院)

AI总结 针对手部运动补全难题,提出先验优先、条件其次框架,先从无结构数据学习通用先验,再引入语义控制,通过分层适配器实现可控性,经评估其在多方面优于基线,还展示了实时推理等特性,适用于动画制作。

详情
AI中文摘要

合成与全身运动和语义标签匹配的手部运动是一项艰巨任务,因其自由度高且缺乏语义标签。为此,我们提出了一种用于身体条件手部运动补全的先验优先、条件其次框架。该框架首先从大规模无结构和无标签运动数据中学习通用的身体 - 手部运动学先验,捕捉全局身体动力学和手部关节之间的内在协调。然后通过在冻结先验之上的轻量级适配引入语义控制,避免为每个控制接口重新学习运动学结构。框架以流式自回归身体 - 手部先验为核心,实时从身体动力学生成连贯、运动学上一致的手部运动。为在有限监督下实现实际可控性,引入语义分层适配器在适当运动学层面注入条件信号。大量评估表明,与端到端条件基线相比,我们的框架提高了运动学合理性、鲁棒性和可控性,尤其在低资源和跨数据集设置中。还展示了实时推理和交互式创作工作流程,突出了其在生产动画管道中的适用性。

英文摘要

Synthesizing hand motion that matches the full body motion and the semantic labels is a difficult task due to their high degrees of freedom and the lack of semantic labels. To cope with this issue, we propose a prior-first, condition-second framework for body-conditioned hand motion completion. Our framework first learns a generic body-hand kinematic prior from large-scale unstructured and unlabeled motion data, capturing the intrinsic coordination between global body dynamics and hand articulation. Semantic control is then introduced through lightweight adaptation on top of the frozen prior, avoiding the need to relearn kinematic structure for each control interface. Our framework centers on a streaming, autoregressive body-hand prior that generates coherent, kinematically consistent hand motion from body dynamics in real time, using structured kinematic modeling to maintain mechanical body-hand coupling. To enable practical controllability under limited supervision, we introduce semantically-layered adapters that inject conditioning signals at appropriate kinematic levels, supporting both self-supervised attribute control and weakly supervised text-driven control with only a few hours of labeled data. Extensive evaluations demonstrate that our framework improves kinematic plausibility, robustness, and controllability compared to end-to-end conditioned baselines, particularly in low-resource and cross-dataset settings. We further showcase real-time inference and an interactive authoring workflow, highlighting the applicability to production animation pipelines. Homepage: https://AIGAnimation.github.io/HandPrior/

URL PDF HTML 收藏
2602.14147 2026-07-08 cs.CV 版本更新

LaViDa-R1: Advancing Reasoning for Unified Multimodal Diffusion Language Models

LaViDa-R1:推进统一多模态扩散语言模型的推理

Shufan Li, Yuchen Zhu, Jiuxiang Gu, Kangning Liu, Zhe Lin, Yongxin Chen, Molei Tao, Aditya Grover, Jason Kuen

机构 * Adobe UCLA(加州大学洛杉矶分校) Georgia Tech(佐治亚理工学院)

AI总结 研究提出多模态通用推理dLLM LaViDa-R1,不同于现有工作,它以统一方式整合多任务,采用新颖统一训练后框架,集成监督微调与多任务强化学习,并运用多种训练技术,在多模态任务上展现强大性能。

Comments 28 pages, 11 figures

详情
AI中文摘要

扩散语言模型(dLLMs)最近成为自回归语言模型的一个有前途的替代方案。最新工作进一步将其扩展到多模态理解和生成任务。在这项工作中,我们提出了LaViDa-R1,一种多模态通用推理dLLM。与通过特定任务强化学习构建推理dLLM的现有工作不同,LaViDa-R1以统一方式整合了多种多模态理解和生成任务。特别是,LaViDa-R1采用了新颖的统一训练后框架,无缝集成了监督微调(SFT)和多任务强化学习(RL)。它采用了几种新颖的训练技术,包括答案强制、树搜索和互补似然估计,以提高有效性和可扩展性。大量实验证明了LaViDa-R1在广泛的多模态任务上的强大性能,包括视觉数学推理、推理密集型基础和图像编辑。

英文摘要

Diffusion language models (dLLMs) recently emerged as a promising alternative to auto-regressive LLMs. The latest works further extended it to multimodal understanding and generation tasks. In this work, we propose LaViDa-R1, a multimodal, general-purpose reasoning dLLM. Unlike existing works that build reasoning dLLMs through task-specific reinforcement learning, LaViDa-R1 incorporates diverse multimodal understanding and generation tasks in a unified manner. In particular, LaViDa-R1 is built with a novel unified post-training framework that seamlessly integrates supervised finetuning (SFT) and multi-task reinforcement learning (RL). It employs several novel training techniques, including answer-forcing, tree search, and complementary likelihood estimation, to enhance effectiveness and scalability. Extensive experiments demonstrate LaViDa-R1's strong performance on a wide range of multimodal tasks, including visual math reasoning, reason-intensive grounding, and image editing.

URL PDF HTML 收藏
2607.05261 2026-07-07 cs.CV 新提交

FlowMark: Mask-Guided Video Watermarking

FlowMark:掩码引导的视频水印方案

Vishal Asnani, Shruti Agarwal, John Collomosse

机构 * Adobe Research(Adobe研究院) DECaDE, University of Surrey(萨里大学DECaDE实验室)

AI总结 该研究提出FlowMark视频水印框架,通过专用掩码预测网络自动确定最优嵌入区域,实现高感知质量下的强鲁棒性,可抵御压缩、时域编辑与主流平台重编码,支持内容溯源等场景。

详情
AI中文摘要

本文提出FlowMark,一种由自动预测的对象掩码引导的视频水印框架。不同于此前需要用户提供掩码引导的基于区域的方法,FlowMark通过专用的Mask Predictor网络学习识别适合水印嵌入的最优区域。该端到端可训练架构融合区域感知编码与噪声增强训练,在保留高感知质量的同时,确保水印对压缩、几何变换和内容变化具备鲁棒性。其内容自适应掩码使水印信号与自然视频动态保持一致,有效消除了感知闪烁。除压缩鲁棒性外,FlowMark在视频原生时域编辑(如帧交换、插入、删除、重采样与插值)以及真实社交媒体分发流水线(如YouTube和Facebook的重编码)下仍能实现可靠的水印恢复。在图像和视频数据集上的实验结果表明,FlowMark可可靠嵌入128比特消息,峰值信噪比(PSNR)最高可达50.08 dB,在内容溯源、时域真实性验证和视频完整性保护场景下表现优异。

英文摘要

We present FlowMark, a video watermarking framework guided by automatically predicted object masks. In contrast to prior region-based approaches that require user-supplied mask guidance, FlowMark learns to identify optimal regions for watermark embedding through a dedicated Mask Predictor network. Our end-to-end trainable architecture combines region-aware encoding with noise-augmented training to ensure robustness against compression, geometric transformations, and content variation, while preserving high perceptual quality. Our content-adaptive masking keeps watermark signals coherent with natural video dynamics, effectively eliminating perceptual flicker. Beyond compression robustness, FlowMark maintains reliable watermark recovery under video-native temporal edits (e.g., frame swap, insertion, deletion, resampling, and interpolation) and real-world social media distribution pipelines (e.g., YouTube and Facebook re-encoding). Experimental results on both image and video datasets show that FlowMark reliably embeds $128$-bit messages with up to $50.08$ dB PSNR, offering strong performance for content provenance, temporal authenticity verification, and video integrity protection.

URL PDF HTML 收藏
2607.04801 2026-07-07 cs.CV 新提交

LILAC: Layer-Wise Independent LoRAs and Cascaded Conditioning for Multi-Concept Customization of Diffusion Models

LILAC:用于扩散模型多概念定制的分层独立LoRAs和级联条件处理

Marian Lupascu, Sebastian Ripa, Mihai Trascau, Mariana-Iuliana Georgescu, Ionut Mironica

机构 * Adobe Research, Romania(罗马尼亚Adobe研究院) Department of Computer Science, University of Bucharest, Romania(罗马尼亚布加勒斯特大学计算机科学系) International Computer High School of Bucharest, Romania(罗马尼亚布加勒斯特国际计算机高中)

AI总结 研究个性化文本到图像扩散模型渲染特定主题的难题,提出LILAC框架,通过分层独立训练低秩适配器并级联条件处理,避免参数干扰,无需联合训练,线性扩展且与主干无关,效果良好。

Comments 19 pages, 8 figures

详情
AI中文摘要

将文本到图像的扩散模型个性化以在连贯图像中渲染多个特定主题仍然具有挑战性。本文提出LILAC框架,在推理时组合独立训练的低秩适配器,每次仅一个适配器激活,避免参数级干扰,无需联合训练,线性扩展且与主干无关,实验效果良好。

英文摘要

Personalizing text-to-image diffusion models to render several specific subjects in a coherent image remains challenging: the model must preserve each subject's identity while keeping the scene spatially and visually coherent. Methods that fuse independently trained concept adapters in a shared weight space (via federated averaging, gradient fusion, or orthogonality constraints) suffer from identity confusion and style bleeding and require joint retraining. In this work, we show that composing concepts as separate image layers, instead of merging their adapters in a shared weight space, avoids parameter-level interference. We introduce LILAC, a framework that composes independently trained low-rank adapters at inference time: each subject is conditioned on the frozen composite of previously placed subjects, with exactly one adapter active at a time, therefore identities never interfere at the parameter level. LILAC composes the adapters without any joint training, scales linearly with the number of concepts, and is backbone-agnostic. Under the Orthogonal Adaptation protocol, LILAC applied on Qwen-Image-Edit reaches an ArcFace detection rate of 0.861, while Orthogonal Adaptation reports 0.745 in its original setting. Adaptation reports 0.745 in its original setting. Code is available at https://github.com/marianlupascu/LILAC.

URL PDF HTML 收藏
2607.04235 2026-07-07 cs.CL 新提交

Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations

化腐朽为神奇:事后重新标记大语言模型智能体轨迹以实现成功示范

Zichao Li, Gang Wu, Zichao Wang, Ruiyi Zhang, Wanrong Zhu, Ryan A. Rossi, Vlad I Morariu, Jihyung Kil

机构 * Mila, McGill University(米拉,麦吉尔大学) Adobe Research(Adobe研究院)

AI总结 针对大语言模型智能体监督瓶颈,利用智能体展开中隐含的成功目标,引入事后监督学习(HSL),通过辅助大语言模型重新标记轨迹及目标并微调,提出两种技术减轻数据次优性,实验证明其优势。

Comments Accepted to ICLR 2026

详情
AI中文摘要

大语言模型智能体在部分可观测的长时环境中运行,获取监督仍是主要瓶颈。我们利用现有训练后方法中被忽视的监督来源——智能体展开中隐含的成功目标来解决此问题。具体而言,我们引入事后监督学习(HSL),其中一个辅助大语言模型审查每个完整轨迹并用智能体实际实现的所有自然语言目标重新标记它。然后,HSL将轨迹与其重新标记的目标配对,并使用这些对进行额外的微调。为了减轻重新标记数据中的次优性,我们为HSL提出了两种学习技术,即无关动作屏蔽和样本重新加权。我们的实验表明,HSL具有灵活性且与现有的训练后管道兼容。它改进了SFT和DPO,在目标空间更多样化的长时任务上有更大的提升。此外,HSL具有样本效率:在ALFWorld上,它仅使用四分之一的真实示范就超过了在完整数据集上训练的基线。

英文摘要

Large language model agents operate in partially observable, long-horizon settings where obtaining supervision remains a major bottleneck. We address this by utilizing a source of supervision overlooked in existing post-training methods: unintended yet successful goals embedded within agent rollouts. Specifically, we introduce Hindsight Supervised Learning (HSL), where an auxiliary LLM reviews each completed trajectory and relabels it with all of the natural-language goals the agent actually achieved. HSL then pairs the trajectory with its relabeled goals and uses these pairs for additional fine-tuning. To mitigate suboptimality in the relabeled data, we propose two learning techniques for HSL, irrelevant-action masking and sample reweighting. Our experiments show that HSL is flexible and compatible with existing post-training pipelines. It improves both SFT and DPO, with larger gains on long-horizon tasks with more diverse goal spaces. Moreover, HSL is sample-efficient: on ALFWorld, it surpasses baselines trained on the full dataset while using only one quarter of the ground-truth demonstrations.

URL PDF HTML 收藏
2607.03675 2026-07-07 cs.LG 新提交

Validation-Induced Shapley Shifts: How Validation Structure Distorts Data Valuation

验证诱导的沙普利值偏移:验证结构如何扭曲数据评估

Yinan Shen, Ziao Yang, Hongfu Liu

机构 * Adobe(奥多比公司)

AI总结 研究发现对验证集的微小改变会致沙普利值分布方向偏移,训练样本沙普利值向零压缩,追踪至噪声诱导邻域重排效应,用KNN-沙普利框架验证,提出策略减轻扭曲。

Comments 12 pages, 8 figures, 1 table

详情
AI中文摘要

沙普利值广泛用于根据训练数据对验证集性能的边际贡献来为其赋值。现有做法常假定训练数据和模型固定时这些值稳定。本文发现即使对验证集做适度改变,如引入噪声,也会导致沙普利分布的方向偏移。添加噪声时,训练样本的沙普利值向零压缩。我们将此追溯到噪声诱导的邻域重排效应:扰动改变了验证和训练样本之间的局部排名顺序,使评估格局扁平化。使用KNN-沙普利框架,我们通过合成数据和真实数据表明这些偏移是一致且可重复的。我们的发现挑战了沙普利稳定性的假设,并揭示了数据评估中一个新的脆弱性轴。我们提出归一化和边界感知验证策略,以减轻这些扭曲,并在机器学习市场中实现更稳健、可解释的评估。

英文摘要

Shapley values are widely used to attribute value to training data based on their marginal contribution to performance on a validation set. Existing practice often assumes these values are stable once the training data and model are fixed. In this work, we uncover a systematic vulnerability: even modest changes to the validation set, such as introducing noises, cause directional shifts in Shapley distributions. As noises are added, Shapley values of training samples compress toward zero. We trace this to a noise-induced neighborhood reshuffling effect: perturbations alter the local rank order between validation and training samples, flattening the valuation landscape. Using the KNN-Shapley framework, we show through synthetic and real data that these shifts are consistent and reproducible. Our findings challenge the assumption of Shapley stability and reveal a new axis of fragility in data valuation. We propose normalization and boundary-aware validation strategies to mitigate these distortions and enable more robust, interpretable valuation in machine learning marketplaces.

URL PDF HTML 收藏
2401.05631 2026-07-07 cs.HC cs.AI cs.CL cs.ET cs.GR

DrawTalking: Building Interactive Worlds by Sketching and Speaking

DrawTalking:通过绘图和说话构建交互世界

Karl Toby Rosenberg, Rubaiat Habib Kazi, Li-Yi Wei, Haijun Xia, Ken Perlin

机构 * New York University(纽约大学) Adobe Research(Adobe研究院) University of California, San Diego(加州大学圣迭戈分校)

AI总结 DrawTalking通过绘图和说话构建交互世界,提供无代码编程能力,适用于创意探索场景,展示其灵活性和应用潜力。

Comments 25 pages, 27 figures; Matching version accepted at UIST 2024

详情
AI中文摘要

我们介绍了DrawTalking,一种通过绘图和说话来构建和控制交互世界的方法。它强调用户控制和灵活性,并提供类似编程的能力,而无需编写代码。我们的原型早期开放性研究显示,其机制具有共鸣性,并适用于许多创意探索性用例,有望启发和影响未来自然界面在创意探索和创作中的研究。

英文摘要

We introduce DrawTalking, an approach to building and controlling interactive worlds by sketching and speaking while telling stories. It emphasizes user control and flexibility, and gives programming-like capability without requiring code. An early open-ended study with our prototype shows that the mechanics resonate and are applicable to many creative-exploratory use cases, with the potential to inspire and inform research in future natural interfaces for creative exploration and authoring.

URL PDF HTML 收藏
2512.09112 2026-07-02 cs.CV 版本更新

GimbalDiffusion: Gravity-Aware Camera Control for Video Generation

GimbalDiffusion:基于重力的摄像机控制用于视频生成

Frédéric Fortier-Chouinard, Yannick Hold-Geoffroy, Valentin Deschaintre, Matheus Gadelha, Jean-François Lalonde

机构 * Université Laval(拉瓦尔大学) Adobe

AI总结 本文提出GimbalDiffusion框架,通过物理坐标系实现摄像机运动的精确控制,引入null-pitch conditioning策略以避免冲突提示内容干扰,并建立新基准评估重力感知摄像机控制视频生成能力。

Comments Project page: https://lvsn.github.io/GimbalDiffusion/

详情
AI中文摘要

近年来,文本到视频生成取得了显著进展,实现了逼真的视频效果,但对摄像机运动和方向的精细控制仍不明确,尤其在极端轨迹(如180度转身或直接仰视/俯视)方面尤为突出。现有方法通常使用相对或模糊的表示来编码摄像机轨迹,限制了精确的几何控制,并对大旋转支持有限。我们引入GimbalDiffusion框架,通过重力作为全局参考,在物理世界坐标系中实现摄像机控制。与基于前一帧的相对运动描述不同,我们的方法在绝对坐标系中定义摄像机轨迹,从而实现对摄像机参数的准确、可解释的控制。利用全景360度视频进行训练,我们覆盖了所有可能的视角,包括超出传统视频数据分布的极端俯仰和翻转组合。为提高摄像机指导,我们引入null-pitch conditioning策略,防止模型在存在冲突提示内容时覆盖摄像机规格(例如,生成草地时摄像机指向天空)。最后,我们提出新的基准以评估重力感知摄像机控制视频生成能力,评估模型生成极端摄像机角度和量化其输入提示纠缠的能力。

英文摘要

Recent progress in text-to-video generation has achieved remarkable realism, yet fine-grained control over camera motion and orientation remains elusive, especially with extreme trajectories (e.g., a 180-degree turnaround, or looking directly up or down). Existing approaches typically encode camera trajectories using relative or ambiguous representations, limiting precise geometric control and offering limited support for large rotations. We introduce GimbalDiffusion, a framework that enables camera control grounded in physical-world coordinates, using gravity as a global reference. Instead of describing motion relative to previous frames, our method defines camera trajectories in an absolute coordinate system, allowing accurate, interpretable control over camera parameters. Using panoramic 360-degree videos for training, we cover the full sphere of possible viewpoints, including combinations of extreme pitch and roll that are out-of-distribution of conventional video data. To improve camera control, we introduce null-pitch conditioning, a strategy that prevents the model from overriding camera specifications in the presence of conflicting prompt content (e.g., generating grass while the camera points toward the sky). Finally, we propose new benchmarks to evaluate gravity-aware camera-controlled video generation, assessing models' ability to generate extreme camera angles and quantify their input prompt entanglement.

URL PDF HTML 收藏