arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

至 收录 71
2606.22804 2026-07-16 cs.CV 版本更新

CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams

CoVStream: 面向长视频流的边缘-云协作理解

Xu Liu, Guikun Chen, Zihao Yan, Kanzhi Wu, Wenguan Wang

机构 * The State Key Lab of Brain-Machine Intelligence, Zhejiang University(浙江大学脑机智能国家重点实验室) vivo Mobile Communication Co., Ltd., Shenzhen, China(深圳 vivo 通信有限公司)

AI总结 提出边缘-云协作框架CoVStream,边缘节点将视频流压缩为视觉特征和语义描述传输至云端,云服务器构建实体图与全局视觉上下文,仅在用户查询时激活大模型,在LVBench上带宽降低87.6%且准确率保持99.2%。

Comments 9 pages

详情
AI中文摘要

长连续视频流是多媒体智能日益关键的驱动力。现有工作通常采用大模型进行采样-编码-推理的方式来处理长视频,但忽略了一个关键的部署事实:视频流通常由计算受限的设备产生。这迫使做出不可行的妥协:云端卸载虽能实现强推理但带来高昂带宽开销,而设备端处理则受限于边缘硬件能力。因此,我们提出CoVStream,首个用于理解长视频流的边缘-云协作框架。边缘节点将原始视频流蒸馏为紧凑的视觉特征和语义描述传输至云端,最小化带宽成本;云服务器将这些数据整合为实体图和全局视觉上下文,仅在用户查询到达时激活重型推理模型。在VideoMME-Long、LVBench和RTV-Bench上的实验表明,CoVStream在LVBench上降低87.6%的带宽使用,同时保留云端基线99.2%的准确率。

英文摘要

Long, continuous video streams are an increasingly critical driver of multimedia intelligence. Existing efforts often handle long videos with a sample-encode-reason approach using large models. However, they overlook a crucial deployment fact: the stream is often produced by computationally constrained devices. This forces an untenable compromise: cloud offloading unlocks strong reasoning but incurs prohibitive bandwidth overhead, while on-device processing remains limited by edge hardware capacity. Therefore, we propose CoVStream, the first edge-cloud collaborative framework for understanding long video streams. The edge node distills raw video streams into compact visual features and semantic captions for transmission to the cloud, minimizing bandwidth costs, while the cloud server integrates this data into an entity graph and global visual context, activating the heavy reasoning model only when a user query arrives. Experiments on VideoMME-Long, LVBench, and RTV-Bench show that CoVStream reduces bandwidth usage by 87.6% while retaining 99.2% of the cloud baseline accuracy on LVBench.

URL PDF HTML 收藏
2508.17434 2026-06-26 cs.CV 版本更新

TinySR: Pruning Diffusion for Real-World Image Super-Resolution

TinySR:剪枝扩散模型用于现实世界图像超分辨率

Linwei Dong, Qingnan Fan, Yuhang Yu, Qi Zhang, Jinwei Chen, Yawei Luo, Changqing Zou

机构 * Zhejiang University(浙江大学) Vivo Mobile Communication Co. Ltd(沃伊通信有限公司) Zhejiang Lab(浙江实验室)

AI总结 本文提出TinySR,一种针对现实世界图像超分辨率的紧凑高效扩散模型,通过动态块激活和扩展-腐蚀策略实现实时性能与高质量输出,相比教师模型TSD-SR在计算成本和参数量上分别减少5.68倍和83%。

详情
AI中文摘要

现实世界图像超分辨率(Real-ISR)旨在从受噪声、模糊和压缩等复杂退化影响的低分辨率输入中恢复高质量图像。最近,扩散模型(DMs)通过利用强大的生成先验来恢复细节,在此领域展现出巨大潜力。然而,其迭代去噪过程带来了高计算开销,对实时应用构成挑战。尽管一步蒸馏方法如OSEDiff和TSD-SR提供更快的推理速度,但它们仍然受限于大、过参数化的模型架构。在本工作中,我们提出TinySR,一种专为Real-ISR设计的紧凑而有效的扩散模型,实现了实时性能同时保持感知质量。我们引入动态块激活和扩展-腐蚀策略以促进深度剪枝更有效的决策。我们通过通道剪枝、注意力移除和轻量级SepConv实现VAE压缩。我们消除了时间相关和提示相关的模块,并执行预缓存技术以进一步加速模型。TinySR显著降低了计算成本和模型大小,相比其教师TSD-SR,计算速度提升达5.68倍,参数量减少83%,同时仍提供高质量的结果。

英文摘要

Real-world image super-resolution (Real-ISR) focuses on recovering high-quality images from low-resolution inputs that suffer from complex degradations like noise, blur, and compression. Recently, diffusion models (DMs) have shown great potential in this area by leveraging strong generative priors to restore fine details. However, their iterative denoising process incurs high computational overhead, posing challenges for real-time applications. Although one-step distillation methods, such as OSEDiff and TSD-SR, offer faster inference, they remain fundamentally constrained by their large, over-parameterized model architectures. In this work, we present TinySR, a compact yet effective diffusion model specifically designed for Real-ISR that achieves real-time performance while maintaining perceptual quality. We introduce a Dynamic Inter-block Activation and an Expansion-Corrosion Strategy to facilitate more effective decision-making in depth pruning. We achieve VAE compression through channel pruning, attention removal and lightweight SepConv. We eliminate time- and prompt-related modules and perform pre-caching techniques to further speed up the model. TinySR significantly reduces computational cost and model size, achieving up to 5.68x speedup and 83% parameter reduction compared to its teacher TSD-SR, while still providing high quality results.

URL PDF HTML 收藏
2606.19195 2026-06-18 cs.CV 新提交

Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance

Moebius: 0.2B轻量级图像修复框架,性能达10B级别

Kangsheng Duan, Ziyang Xu, Wenyu Liu, Xiaohu Ruan, Xiaoxin Chen, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) VIVO AI Lab(VIVO人工智能实验室)

AI总结 提出Moebius轻量级图像修复框架,通过局部-λ混合交互模块和自适应多粒度蒸馏策略,以0.22B参数实现与10B级模型FLUX.1-Fill-Dev相当甚至更优的生成质量,推理速度提升15倍以上。

详情
AI中文摘要

尽管10B级别的工业基础模型推动了图像修复的边界,但其高昂的计算成本严重阻碍了实际部署。构建高度优化的任务特定专家模型是一个有前景的解决方案,然而极端的结构压缩不可避免地引发了严重的表示瓶颈。为解决这一问题,我们提出了Moebius,一个高效的轻量级修复框架。我们通过引入局部-λ混合交互($L\lambda MI$)模块系统地重构了扩散主干。该模块由局部-λ和交互-λ子模块组成,巧妙地将空间上下文和全局语义先验总结为固定大小的线性矩阵,在保留复杂潜在交互的同时大幅减少参数。此外,为了释放这种高度紧凑架构的全部表示能力,我们将其与自适应多粒度蒸馏策略协同配对。该策略严格在潜在空间内操作以避免昂贵的像素空间解码,动态平衡多个基于梯度的损失以实现高保真对齐。在自然和肖像基准上的大量实验表明,这种最优协同使Moebius能够媲美甚至超越10B级工业通用模型FLUX.1-Fill-Dev的生成质量。值得注意的是,Moebius仅使用不到2%的参数(0.22B vs. 11.9B)就实现了这一点,同时总推理时间加速超过15倍,为高保真修复设立了新的效率标准。项目页面见此https URL。

英文摘要

While 10B-level industrial foundation models have pushed the boundaries of image inpainting, their prohibitive computational costs severely hinder practical deployment. Constructing a highly optimized task-specific specialist offers a promising solution; however, extreme structural compression inevitably triggers a severe representation bottleneck. To conquer this, we propose Moebius, a highly efficient lightweight inpainting framework. We systematically reconstruct the diffusion backbone by introducing the Local-$λ$ Mix Interaction ($LλMI$) block. Comprising Local-$λ$ and Interactive-$λ$ modules, it elegantly summarizes spatial contexts and global semantic priors into fixed-size linear matrices, preserving complex latent interactions while drastically shedding parameters. Furthermore, to unlock the full representational capacity of this highly compact architecture, we synergistically pair it with an adaptive multi-granularity distillation strategy. Operating strictly within the latent space to avoid expensive pixel-space decoding, this strategy dynamically balances multiple gradient-based losses to achieve high-fidelity alignment. Extensive experiments across natural and portrait benchmarks demonstrate that this optimal synergy enables Moebius to rival or even surpass the generation quality of the 10B-level industrial generalist FLUX.1-Fill-Dev. Remarkably, Moebius achieves this using less than 2\% of the parameters (0.22B vs. 11.9B) while delivering a $>15\times$ acceleration in total inference time, setting a new efficiency standard for high-fidelity inpainting. Project page at https://hustvl.github.io/Moebius.

URL PDF HTML 收藏
2605.30039 2026-06-01 cs.AI

Domain-Specific Data Synthesis for LLMs via Minimal Sufficient Representation Learning

基于最小充分表示学习的大语言模型领域特定数据合成

Tong Ye, Hang Yu, Tengfei Ma, Xuhong Zhang, Jianguo Li, Peng Di, Peiyu Liu, Jianwei Yin, Wenhai Wang

机构 * vivo AI Lab(vivo人工智能实验室) Ant Group(蚂蚁集团) Zhejiang University(浙江大学)

AI总结 提出DOMINO框架,通过对比解耦学习最小充分领域表示,指导生成领域对齐的合成数据,在隐式领域定义下提升微调性能。

Comments Accepted by KDD 2026

详情
AI中文摘要

大语言模型在通用能力上取得了显著进展,并可通过在领域特定数据上微调在特定领域实现强性能。然而,获取目标领域的高质量数据仍是一个重大挑战。现有数据合成方法遵循演绎范式,严重依赖自然语言表达的显式领域描述和精心设计的提示工程,限制了其在领域难以描述或正式表述的现实场景中的适用性。在这项工作中,我们通过归纳范式处理未被充分探索的领域特定数据合成问题,其中目标领域仅通过一组参考示例定义,特别是在领域特征难以用自然语言表述时。我们提出了一种新颖框架DOMINO,它从参考样本中学习最小充分的领域表示,并利用它来指导生成领域对齐的合成数据。DOMINO将提示调优与对比解耦目标相结合,以分离领域级模式与样本特定噪声,在保留核心领域特征的同时缓解过拟合。理论上,我们证明DOMINO扩展了合成数据分布的支持集,确保了更大的多样性。在隐式领域定义的具有挑战性的编码基准上,对DOMINO合成的数据进行微调,在强大的指令调优基线上将Pass@1准确率提高了高达4.63%,证明了其有效性和鲁棒性。这项工作为领域特定数据合成建立了一种新范式,无需手动提示设计或自然语言领域规范即可实现实用且可扩展的领域适应。

英文摘要

Large Language Models have demonstrated remarkable progress in general-purpose capabilities and can achieve strong performance in specific domains through fine-tuning on domain-specific data. However, acquiring high-quality data for target domains remains a significant challenge. Existing data synthesis approaches follow a deductive paradigm, heavily relying on explicit domain descriptions expressed in natural language and careful prompt engineering, limiting their applicability in real-world scenarios where domains are difficult to describe or formally articulate. In this work, we tackle the underexplored problem of domain-specific data synthesis through an inductive paradigm, where the target domain is defined only through a set of reference examples, particularly when domain characteristics are difficult to articulate in natural language. We propose a novel framework, DOMINO, that learns a minimal sufficient domain representation from reference samples and leverages it to guide the generation of domain-aligned synthetic data. DOMINO integrates prompt tuning with a contrastive disentanglement objective to separate domain-level patterns from sample-specific noise, mitigating overfitting while preserving core domain characteristics. Theoretically, we prove that DOMINO expands the support of the synthetic data distribution, ensuring greater diversity. Empirically, on challenging coding benchmarks where domain definitions are implicit, fine-tuning on data synthesized by DOMINO improves Pass@1 accuracy by up to 4.63\% over strong, instruction-tuned backbones, demonstrating its effectiveness and robustness. This work establishes a new paradigm for domain-specific data synthesis, enabling practical and scalable domain adaptation without manual prompt design or natural language domain specifications.

URL PDF HTML 收藏
2605.25447 2026-05-26 cs.CL

GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation

GeoSVG-RL:面向布局约束的文本到SVG图表生成的几何感知强化学习

Sifan Li, Yujun Cai, Hongkai Chen, Yiwei Wang

机构 * University of California, Merced(加州大学梅尔德分校) The University of Queensland(昆士兰大学) vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司)

AI总结 提出GeoSVG-RL框架,通过强化学习优化策略,利用几何反馈奖励(渲染有效性、画布适配、锚点放置、文本包含、图一致性和代码整洁性)解决文本到SVG图表生成中的结构脆弱性问题,显著提升箭头锚点精度和文本框内率。

详情
AI中文摘要

生成结构化、可编辑的图表对当代大型语言模型来说仍然是一个重大挑战,尽管它们在通用向量代码生成方面表现出色。主要困难在于输出的结构脆弱性;微小的错误,如未对齐的连接器端点、文本标签与边框重叠或复杂布局超出画布边界,都会使生成的SVG文件在专业应用中无法使用。为了解决这些问题,我们引入了GeoSVG-RL,一个专门为布局约束的文本到SVG生成设计的强化学习框架。与仅依赖于最大化令牌级可能性的标准训练目标不同,我们的方法针对明确的、可执行的几何反馈优化策略。模型首先生成一个结构化的布局计划,作为后续SVG代码生成的几何契约。然后通过浏览器支持的验证器渲染该代码,从而在六个关键维度上计算细粒度奖励:渲染有效性、画布适配、精确锚点放置、文本包含、图一致性和代码整洁性。我们利用组相对策略优化(GRPO)来优化模型,每个提示采样多个候选,以便基于相对质量进行更新。从合成数据上的监督预热阶段开始,GeoSVG-RL在结构可靠性方面取得了显著提升,特别是在箭头锚点精度和文本框内率方面。定量评估表明,我们的方法在局部几何精度和图连通性保持方面持续优于当前最先进的系统,为自动化且可靠的技术插图提供了一条稳健的路径。

英文摘要

Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-purpose vector code generation. The primary difficulty lies in the structural fragility of the output; minor errors such as misaligned connector endpoints, text labels overlapping borders, or complex layouts drifting beyond the canvas boundaries render the resulting SVG files functionally unusable for professional applications. To address these issues, we introduce GeoSVG-RL, a specialized reinforcement learning framework designed for layout-constrained text-to-SVG generation. Unlike standard training objectives that rely solely on maximizing token-level likelihood, our approach optimizes the policy against explicit, executable geometric feedback. The model first produces a structured layout plan that serves as a geometric contract for the subsequent generation of the SVG code. This code is then rendered through a browser-backed verifier, enabling the calculation of fine-grained rewards across six critical dimensions: rendering validity, canvas fitting, precise anchor placement, text containment, graph consistency, and code cleanliness. We utilize Group Relative Policy Optimization (GRPO) to refine the model, sampling multiple candidates per prompt to facilitate updates based on relative quality. Starting from a supervised warm-start phase on synthetic data, GeoSVG-RL achieves substantial gains in structural reliability, particularly in arrow-anchor accuracy and text-in-box rates. Quantitative evaluations demonstrate that our method consistently outperforms current state-of-the-art systems in local geometric precision and the preservation of graph connectivity, providing a robust pathway toward automated yet reliable technical illustration.

URL PDF HTML 收藏
2603.08155 2026-05-21 cs.LG

C$^2$FG: Control Classifier-Free Guidance via Score Discrepancy Analysis

C$^2$FG: 通过分数差异分析实现控制分类器无关引导

Jiayang Gao, Tianyi Zheng, Jiayang Zou, Fengxiang Yang, Shice Liu, Luyao Fan, Zheyu Zhang, Hao Zhang, Jinwei Chen, Peng-Tao Jiang, Bo Li, Jia Wang

机构 * Shanghai Jiao Tong University(上海交通大学) vivo BlueImage Lab(vivo 蓝影实验室) vivo Mobile Communication Co., Ltd.(vivo 通信有限公司)

AI总结 本文提出C$^2$FG,一种基于分数差异分析的控制分类器无关引导方法,通过严格理论分析建立了条件分布与无条件分布在不同时间步的分数差异上界,从而为时间依赖引导提供了理论基础,并通过实验验证了其在多种生成任务中的有效性。

Comments Accepted to CVPR 2026 (Highlight)

详情
AI中文摘要

分类器无关引导(CFG)是现代条件扩散模型的核心,但其依赖于固定或启发式动态引导权重,主要基于经验,忽略了扩散过程的内在动态。本文对分类器无关引导进行了严格的理论分析。具体而言,我们基于扩散过程建立了条件分布与无条件分布在不同时间步的分数差异的严格上界。这一发现解释了固定权重策略的局限性,并为时间依赖引导建立了原理基础。受此启发,我们引入了控制分类器无关引导(C$^2$FG),一种新颖的、无需训练且可直接使用的插件方法,通过指数衰减控制函数将引导强度与扩散动态对齐。大量实验表明,C$^2$FG在多种生成任务中均有效且具有广泛的应用性,同时与现有策略具有正交性。

英文摘要

Classifier-Free Guidance (CFG) is a cornerstone of modern conditional diffusion models, yet its reliance on the fixed or heuristic dynamic guidance weight is predominantly empirical and overlooks the inherent dynamics of the diffusion process. In this paper, we provide a rigorous theoretical analysis of the Classifier-Free Guidance. Specifically, we establish strict upper bounds on the score discrepancy between conditional and unconditional distributions at different timesteps based on the diffusion process. This finding explains the limitations of fixed-weight strategies and establishes a principled foundation for time-dependent guidance. Motivated by this insight, we introduce \textbf{Control Classifier-Free Guidance (C$^2$FG)}, a novel, training-free, and plug-in method that aligns the guidance strength with the diffusion dynamics via an exponential decay control function. Extensive experiments demonstrate that C$^2$FG is effective and broadly applicable across diverse generative tasks, while also exhibiting orthogonality to existing strategies.

URL PDF HTML 收藏
2605.17470 2026-05-20 cs.CV cs.MM eess.IV

EchoSR: Efficient Context Harnessing for Lightweight Image Super-Resolution

EchoSR: 为轻量图像超分辨率实现高效的上下文利用

Hanli Zhao, Binhao Wang, Shihao Zhao, Tao Wang, Kaihao Zhang, Wanglong Lu

机构 * College of Computer Science and Artificial Intelligence, Wenzhou University, Wenzhou 325000, China(温州大学计算机科学与人工智能学院) vivo BlueImage Lab, vivo Mobile Communication Co., Ltd, Shanghai 200100, China(vivo蓝影实验室,vivo移动通信有限公司,上海200100,中国) College of Engineering and Computer Science, Australian National University, Canberra, Australia(工程与计算机科学学院,澳大利亚国立大学,堪培拉,澳大利亚) The AI/Analytics Team, Nasdaq, St. John’s, Canada(AI/分析团队,纳斯达克,圣约翰,加拿大)

AI总结 本文提出EchoSR框架,通过统一多尺度感受野建模和层次化上下文融合,提升了轻量图像超分辨率的效率和效果,同时在多个基准上优于现有方法,并实现了约两倍的速度提升。

Comments Accepted by Information Fusion; 20 pages, 17 figures

详情
AI中文摘要

图像超分辨率(SR)旨在从低分辨率(LR)输入中重建高质量、高分辨率(HR)图像,并在各种下游应用中发挥关键作用。尽管近年来取得了进展,但平衡重建保真度和计算效率仍然是一个根本性挑战,尤其是在资源受限的场景中。虽然现有轻量方法试图扩展感受野,但许多方法要么导致显著的计算开销,要么简单地扩大内核大小,或缺乏机制进行一致的多尺度整合,限制了它们的整体效果和可扩展性。为了解决这些限制,我们提出了EchoSR,一个高效的上下文利用框架,用于轻量图像超分辨率,它统一了多尺度感受野建模和层次化上下文融合。EchoSR通过一种高效的上下文利用策略将特征学习解耦为分离的局部、多尺度和全局建模阶段,并进一步通过跨尺度重叠融合机制促进无缝的跨尺度整合。广泛的实验表明,EchoSR在多个基准上一致优于现有最先进的轻量超分辨率方法,同时也实现了更快的速度(约2倍)。源代码可在https://github.com/funnyWang-Echoes/EchoSR上获得。

英文摘要

Image super-resolution (SR) aims to reconstruct high-quality, high-resolution (HR) images from low-resolution (LR) inputs and plays a critical role in various downstream applications. Despite recent advancements, balancing reconstruction fidelity and computational efficiency remains a fundamental challenge, particularly in resource-constrained scenarios. While existing lightweight methods attempt to expand receptive fields, many of them either incur substantial computational overhead, naively scale up kernel sizes, or lack mechanisms for coherent multi-scale integration, limiting their overall effectiveness and scalability. To address these limitations, we propose EchoSR, an efficient context-harnessing framework for lightweight image super-resolution, which unifies multi-scale receptive field modeling and hierarchical context fusion. EchoSR decouples feature learning into disentangled local, multi-scale, and global modeling stages through an efficient context-harnessing strategy, and further promotes seamless cross-scale integration via a cross-scale overlapping fusion mechanism. Extensive experiments have shown that EchoSR consistently outperforms state-of-the-art lightweight super-resolution methods across multiple benchmarks, while also achieving a faster speed $(\sim 2\times)$. The source code is available at https://github.com/funnyWang-Echoes/EchoSR.

URL PDF HTML 收藏
2605.15908 2026-05-18 cs.CV cs.AI

RaPD: Resolution-Agnostic Pixel Diffusion via Semantics-Enriched Implicit Representations

RaPD:通过语义增强的隐式表示实现分辨率无关的像素扩散

Yanhao Ge, Shanyan Guan, Weihao Wang, Ying Tai, Mingyu You

机构 * College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院) vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司) Nanjing University(南京大学)

AI总结 RaPD通过语义表示引导和坐标查询注意力渲染器,在连续神经图像场的潜在空间中实现分辨率无关的像素扩散,解决了重建与生成之间的差距,提升了生成质量和分辨率扩展能力。

详情
AI中文摘要

自然图像是连续的,但大多数生成模型在离散网格上合成图像,限制了分辨率灵活生成。连续神经场使分辨率无关渲染成为可能,但先前方法仅在解码阶段引入连续性作为插值模块,使生成的潜在空间离散化且偏向重建。我们提出RaPD(分辨率无关像素扩散),在连续神经图像场(NIF)潜在空间中进行扩散。RaPD通过语义表示引导实现生成意识的潜在学习,并通过坐标查询注意力渲染器实现坐标条件化的、尺度感知的渲染。通过仅改变查询坐标,单个去噪潜在态可以在任意分辨率下渲染,保持扩散成本不变。实验表明生成质量和分辨率扩展能力均优于现有方法。

英文摘要

Natural images are continuous, yet most generative models synthesize them on discrete grids, limiting resolution-flexible generation. Continuous neural fields enable resolution-free rendering, but prior methods introduce continuity only at the decoding stage as an interpolation module, leaving the generative latent space discretized and reconstruction-oriented. We propose RaPD (Resolution-agnostic Pixel Diffusion), which performs diffusion in a continuous Neural Image Field (NIF) latent space. RaPD bridges this reconstruction-generation gap with Semantic Representation Guidance for generation-aware latent learning and a Coordinate-Queried Attention Renderer for coordinate-conditioned, scale-aware rendering. A single denoised latent can be rendered at arbitrary resolutions by changing only the query coordinates, keeping diffusion cost fixed. Experiments demonstrate superior generation quality and resolution scalability.

URL PDF HTML 收藏
2605.14938 2026-05-15 cs.LG cs.CV

Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models

Octopus:无历史梯度正交化用于多模态大语言模型中的持续学习

Yuehao Liu, Shanyan Guan, Weijia Zhang, Xuanming Shang, Yanhao Ge, Wei Li, Chao Ma

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能大模型关键实验室,人工智能研究院,上海交通大学) vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司)

AI总结 Octopus提出一种基于无历史梯度正交化的方法,通过两阶段微调策略平衡持续学习中的可塑性与稳定性,实验表明其在UCIT数据集上性能优于现有最佳方法。

详情
AI中文摘要

Octopus提出一种基于无历史梯度正交化的方法,通过两阶段微调策略平衡持续学习中的可塑性与稳定性,实验表明其在UCIT数据集上性能优于现有最佳方法。

英文摘要

Continual learning in multimodal large language models (MLLMs) aims to sequentially acquire knowledge while mitigating catastrophic forgetting, yet existing methods face inherent limitations: architecture-based approaches incur additional computational overhead and often generalize poorly to new tasks, rehearsal-based methods rely on storing historical data, raising privacy and storage concerns, and conventional regularization-based strategies alone are insufficient to fully prevent parameter interference. We propose Octopus, a two-stage continual learning framework based on History-Free Gradient Orthogonalization (HiFGO), which enforces gradient-level orthogonality without historical task data. Our proposed two-stage finetuning strategy decouples task adaptation from regularization, achieving a principled balance between plasticity and stability. Experiments on UCIT show that Octopus establishes state-of-the-art performance, surpassing prior SOTA by 2.14% and 6.82% in terms of Avg and Last.

URL PDF HTML 收藏
2605.14821 2026-05-15 cs.CV

HDRFace: Rethinking Face Restoration with High-Dimensional Representation

HDRFace: 重新思考高维表示下的面部修复

Zirui Wang, Xianhui Lin, Yi Dong, Bo Wei, Gangjian Zhang, Siteng Ma, Zebiao Zheng, Xing Liu, Hong Gu, Minjing Dong

机构 * City University of Hong Kong(香港城市大学) vivo BlueImage Lab, vivo Mobile Communication Co., Ltd(vivo蓝影实验室,vivo移动通信有限公司)

AI总结 本文提出HDRFace框架,通过注入语义丰富的先验知识提升面部修复效果,采用高维特征编码器提取细粒度面部表示,并引入SDFM机制平衡结构一致性和细节保真度。

详情
AI中文摘要

在复杂的退化条件下,面部修复仍是一个病态的逆问题,由于信息丢失严重。尽管扩散模型受益于强大的生成先验,但大多数方法仅基于低质量输入进行条件化,难以在重退化下恢复身份关键细节。本文提出HDRFace,一种高维表示条件化的面部修复框架,将语义丰富的先验注入条件流而不修改生成主干。我们的流程首先使用现成的修复器获得结构可靠的中间修复结果,然后使用预训练的高维特征编码器从低质量输入和中间结果中提取细粒度面部表示,并将其作为额外条件用于生成。我们进一步引入SDFM,一种结构-细节感知的自适应融合机制,在结构建模中强调全局约束,在细节合成中增强表示引导,平衡结构一致性与细节保真度。为了验证我们方法的泛化能力,我们在两个生成模型SD V2.1-base和Qwen-Image上实现所提出的框架,并一致观察到在不同架构下均获得稳定且一致的性能提升。

英文摘要

Face restoration under complex degradations still remains an ill-posed inverse problem due to severe information loss. Although diffusion models benefit from strong generative priors, most methods still condition only on low-quality inputs, making it difficult to recover identity-critical details under heavy degradations. In this work, we propose HDRFace, a High-Dimensional Representation conditioned Face restoration framework that injects semantically rich priors into the conditional flow without modifying the generative backbone. Our pipeline first obtains a structurally reliable intermediate restoration with an off-the-shelf restorer, then uses a pretrained high-dimensional feature encoder to extract fine-grained facial representations from both the low-quality input and the intermediate result, and injects them as additional conditions for generation. We further introduce SDFM, a Structure-Detail aware adaptive Fusion Mechanism that emphasizes global constraints during structure modeling and strengthens representation guidance during detail synthesis, balancing structural consistency and detail fidelity. To validate the generalization ability of our method, we implement the proposed framework on two generative models, SD V2.1-base and Qwen-Image, and consistently observe stable and coherent performance gains across different architectures.

URL PDF HTML 收藏
2605.11475 2026-05-13 cs.CV

Deep Probabilistic Unfolding for Quantized Compressive Sensing

深度概率展开用于量化压缩感知

Gang Qu, Ping Wang, Siming Zheng, Xin Yuan

机构 * Westlake University, School of Engineering, Hangzhou, Zhejiang, China(西湖大学工程学院,杭州,浙江,中国) Vivo Mobile Communication Co., Ltd., Hangzhou, Zhejiang, China(Vivo移动通信有限公司,杭州,浙江,中国)

AI总结 本文提出深度概率展开模型,通过展开框架提升重建精度和效率,采用闭式概率梯度投影替代传统L2投影,设计双域Mamba模块融合多尺度特征,实验证明其在量化压缩感知中的优越性能。

详情
AI中文摘要

我们提出了一种深度概率展开模型,以解决经典量化压缩感知问题,利用展开框架提高重建精度和效率。与以往使用L2投影的方法不同,我们推导出闭式且数值稳定的似然梯度投影,使模型能够尊重真实的量化物理,将硬量化约束转化为软概率指导。此外,设计了一个高效的双域Mamba模块,专门用于动态捕捉和融合多尺度局部和全局特征,确保远距离但相关的区域之间的交互。广泛实验表明,所提方法在先前工作中表现出最先进的性能,能够推动量化压缩感知在现实生活中的应用。

英文摘要

We propose a deep probabilistic unfolding model to address the classical quantized compressive sensing problem that leverages an unfolding framework to enhance the reconstruction accuracy and efficiency. Unlike previous unfolding methods that apply L2 projection to measurements, we derive a closed-form, numerically stable likelihood gradient projection, which allows the model to respect the true quantization physics, turning the hard quantization constraint into a soft probabilistic guidance. Furthermore, an efficient, dual-domain Mamba module is specifically designed to dynamically capture and fuse the multi-scale local and global features, ensuring the interactions between the distant but correlated regions. Extensive experiments demonstrate the state-of-the-art performance of the proposed method over previous works, which is capable of promoting the application of quantized compressive sensing in real life.

URL PDF HTML 收藏
2605.07429 2026-05-12 cs.CV

Towards Photorealistic and Efficient Bokeh Rendering via Diffusion Framework

通过扩散框架实现逼真且高效的bokeh渲染

Linxiao Shi, Siming Zheng, Zerong Wang, Hao Zhang, Jinwei Chen, Bo Li, Shifeng Chen, Peng-Tao Jiang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) vivo BlueImage Lab, vivo Mobile Communication Co., Ltd.(vivo BlueImage实验室,vivo移动通信有限公司) Shenzhen University of Advanced Technology(深圳大学)

AI总结 本文提出MagicBokeh框架,通过替代训练策略和聚焦感知的掩码注意力机制,联合优化bokeh渲染与超分辨率,提升可控性和视觉保真度,并引入降质感知深度模块以提高低质量输入的深度估计准确性。

Comments Accepted by CVPR 2026

详情
AI中文摘要

现有移动设备受限于紧凑光学设计,如小孔径,难以产生自然、光学真实的bokeh效果。尽管近期基于学习的方法已展现 promising 结果,但它们在高数字变焦拍摄的照片中仍面临分辨率降低和细节丢失的问题。一种简单方案是增强图像质量后再应用bokeh渲染,但这种两阶段流程降低了效率并引入了不必要的误差累积。为克服这些限制,我们提出MagicBokeh,一种统一的扩散框架,用于高质量且高效的bokeh渲染。通过替代训练策略和聚焦感知的掩码注意力机制,我们的方法联合优化bokeh渲染和超分辨率,显著提高了可控性和视觉保真度。此外,我们引入了降质感知深度模块,以从低质量输入中实现更准确的深度估计。实验结果表明,MagicBokeh能够高效地生成逼真bokeh效果,特别是在真实世界的低分辨率图像上,为未来bokeh渲染的发展铺平了道路。我们的代码和模型可在https://github.com/vivoCameraResearch/MagicBokeh上获得。

英文摘要

Existing mobile devices are constrained by compact optical designs, such as small apertures, which make it difficult to produce natural, optically realistic bokeh effects. Although recent learning-based methods have shown promising results, they still struggle with photos captured under high digital zoom levels, which often suffer from reduced resolution and loss of fine details. A naive solution is to enhance image quality before applying bokeh rendering, yet this two-stage pipeline reduces efficiency and introduces unnecessary error accumulation. To overcome these limitations, we propose MagicBokeh, a unified diffusion-based framework designed for high-quality and efficient bokeh rendering. Through an alternative training strategy and a focus-aware masked attention mechanism, our method jointly optimizes bokeh rendering and super-resolution, substantially improving both controllability and visual fidelity. Furthermore, we introduce degradation-aware depth module to enable more accurate depth estimation from low-quality inputs. Experimental results demonstrate that MagicBokeh efficiently produces photorealistic bokeh effects, particularly on real-world low-resolution images, paving the way for future advancements in bokeh rendering. Our code and models are available at https://github.com/vivoCameraResearch/MagicBokeh.

URL PDF HTML 收藏
2605.07457 2026-05-11 cs.CV

EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

EditRefiner:一种对齐人类的代理框架用于图像编辑细化

Zitong Xu, Huiyu Duan, Yifei Nie, Mingda Du, Sijing Wu, Xiongkuo Min, Tianyi Zheng, Jian Zhang, Shusong Xu, Jinwei Chen, Bo Li, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Vivo Mobile Communication Co., Ltd(Vivo移动通信有限公司) University of Electronic Science and Technology of China(电子科学与技术大学)

AI总结 本文提出EditRefiner,一种对齐人类的代理框架,通过人类反馈数据集改进图像编辑的细粒度问题,通过感知-推理-行动-评估循环提升编辑质量。

详情
AI中文摘要

最近的文本引导图像编辑(TIE)模型取得了显著进展,但编辑后的图像仍经常出现细粒度问题,如不自然的对象、光照不匹配和意外变化。现有的细化方法要么依赖于昂贵的迭代再生,要么使用视觉语言模型(VLMs)具有弱空间定位,往往导致语义漂移和不可靠的局部修正。为了解决这些限制,我们首先构建了EditFHF-15K数据集,包含15K张来自12个TIE模型的图像,涵盖43种编辑任务,60K个标注的瑕疵区域和80K个编辑失败区域,每个区域都配有文本推理,以及45K个平均意见分数(MOSs)评估感知质量、指令遵循和视觉一致性。基于EditFHF-15K,我们提出了EditRefiner,一种分层、可解释且对齐人类的代理框架,将后编辑修正重新表述为人类般的感知-推理-行动-评估循环。具体来说,我们引入:(1)一个感知代理,检测瑕疵和编辑失败的上下文显著图;(2)一个推理代理,解释这些感知线索以执行对齐人类的诊断推理;(3)一个行动代理,使用推理输出计划和执行局部重编辑;(4)一个评估代理,评估重编辑的图像并指导行动代理是否需要进一步细化。广泛的实验表明,EditRefiner在畸变定位、诊断准确率和人类感知对齐方面始终优于最先进的方法,建立了自我纠正和感知可靠的图像编辑新范式。代码可在https://github.com/IntMeGroup/EditRefiner获取。

英文摘要

Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as unnatural objects, lighting mismatch, and unexpected changes. Existing refinement approaches either rely on costly iterative regeneration or employ vision-language models (VLMs) with weak spatial grounding, often resulting in semantic drift and unreliable local corrections. To address these limitations, we first construct EditFHF-15K, a dataset of fine-grained human feedback for edited images, comprising (1) 15K images from 12 TIE models spanning 43 editing tasks, (2) 60K annotated artifact regions and 80K editing failure regions, each accompanied by textual reasoning, and (3) 45K mean opinion scores (MOSs) assessing perceptual quality, instruction following, and visual consistency. Based on EditFHF-15K, we propose EditRefiner, a hierarchical, interpretable, and human-aligned agentic framework that reformulates post-editing correction as a human-like perception-reasoning-action-evaluation loop. Specifically, we introduce: (1) a perception agent that detects contextual saliency maps of artifacts and editing failures, (2) a reasoning agent that interprets these perceptual cues to perform human-aligned diagnostic inference, (3) an action agent that uses the reasoning output to plan and execute localized re-editing, and (4) an evaluation agent that assesses the re-edited image and guides the action agent on whether further refinements are required. Extensive experiments demonstrate that EditRefiner consistently outperforms state-of-the-art methods in distortion localization, diagnose accuracy and human perception alignment, establishing a new paradigm for self-corrective and perceptually reliable image editing. The code is available at https://github.com/IntMeGroup/EditRefiner.

URL PDF HTML 收藏
2605.06708 2026-05-11 cs.CV cs.AI

Visual Text Compression as Measure Transport

视觉文本压缩作为度量传输

Lv Tang, Tianyi Zheng, Yang Liu, Bo Li, Xingyu Li

机构 * University of Alberta(阿尔伯塔大学) vivo Mobile Communication Co., Ltd(vivo移动通信有限公司) Tsinghua University(清华大学)

AI总结 本文通过度量传输理论分析视觉文本压缩,提出无标签路由准则和传输感知聚焦机制,提升压缩效率并优化下游任务表现。

详情
AI中文摘要

视觉文本压缩(VTC)通过将文本渲染为图像并用视觉-语言模型重新编码,实现长上下文处理的高效性,但其压缩比并不直接转化为下游任务的实用性。本文通过度量传输理论,将文本和视觉标记视为经验概率测度,揭示ViT补丁编码器诱导的推前映射的传输成本,包含精度成本和覆盖成本。该方法提出无标签路由准则和传输感知聚焦机制,在24个NLP数据集上,无标签规则在17个数据集上达到Oracle水平,平均任务得分提升3.3%且平均tokens减少10.3%。

英文摘要

Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language model, often producing $3$--$20\times$ fewer decoder tokens than subword tokenization. Yet token savings do not translate predictably into downstream utility: on some tasks the visual path matches or exceeds the text path, on others it collapses, and the compression ratio itself does not predict which regime will occur. The missing quantity is therefore not another summary of efficiency, but a principled measure of task-relevant information loss induced by visual encoding. We address this problem by formulating VTC in the language of measure transport. Treating text and visual tokens as empirical probability measures, we show that the ViT patch encoder induces a push-forward map whose transport cost decomposes into a precision cost from within-patch aggregation and a coverage cost from cross-patch fragmentation. Both terms are estimable from downstream-label-free probes. This formulation yields two operational consequences: a downstream-label-free routing criterion that selects whether to use the visual path for a given input or benchmark instance, and a transport-informed foveation mechanism that re-encodes high-cost regions at higher resolution. Across $24$ NLP datasets at Qwen3-4B, our label-free rule matches the per-dataset oracle on $17/24$ datasets ($70.8\%$), and improves the average task score by $+3.3\%$ with $-10.3\%$ average tokens relative to a pure-LLM.

URL PDF HTML 收藏
2604.22558 2026-04-27 cs.LG cs.AI

SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning

SOLAR-RL:半在线长 horizon 分配强化学习

Jichao Wang, Liuyang Bian, Yufeng Zhou, Han Xiao, Yue Pan, Guozhi Wang, Hao Wang, Zhaoxiong Wang, Yafei Wen, Xiaoxin Chen, Shuai Ren, Lingfang Zeng

机构 * vivo AI Lab(vivo人工智能实验室) Zhejiang Lab(浙江实验室) CUHK MMLab(香港大学多模态实验室) Hubei University(湖北省大学) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院)

AI总结 SOLAR-RL通过整合全局轨迹信息到离线学习中,提升GUI任务的长horizon完成率和鲁棒性,提供高效样本利用的自主导航方案。

Comments 14 pages, 11 figures. Accepted to Findings of the Association for Computational Linguistics: ACL 2026

详情
AI中文摘要

SOLAR-RL通过整合全局轨迹信息到离线学习中,提升GUI任务的长horizon完成率和鲁棒性,提供高效样本利用的自主导航方案。

英文摘要

As Multimodal Large Language Models (MLLMs) mature, GUI agents are evolving from static interactions to complex navigation. While Reinforcement Learning (RL) has emerged as a promising paradigm for training MLLM agents on dynamic GUI tasks, its effective application faces a dilemma. Standard Offline RL often relies on static step-level data, neglecting global trajectory semantics such as task completion and execution quality. Conversely, Online RL captures the long-term dynamics but suffers from high interaction costs and potential environmental instability. To bridge this gap, we propose SOLAR-RL (Semi-Online Long-horizon Assignment Reinforcement Learning). Instead of relying solely on expensive online interactions, our framework integrates global trajectory insights directly into the offline learning process. Specifically, we reconstruct diverse rollout candidates from static data, detect the first failure point using per-step validity signals, and retroactively assign dense step-level rewards with target-aligned shaping to reflect trajectory-level execution quality, effectively simulating online feedback without interaction costs. Extensive experiments demonstrate that SOLAR-RL significantly improves long-horizon task completion rates and robustness compared to strong baselines, offering a sample-efficient solution for autonomous GUI navigation.

URL PDF HTML 收藏
2604.19587 2026-04-22 cs.CV

SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing

SmartPhotoCrafter: 统一推理、生成与优化用于自动摄影图像编辑

Ying Zeng, Miaosen Luo, Guangyuan Li, Yang Yang, Ruiyang Fan, Linxiao Shi, Qirui Yang, Jian Zhang, Chengcheng Liu, Siming Zheng, Jinwei Chen, Bo Li, Peng-Tao Jiang

机构 * vivo BlueImage Lab, vivo Mobile Communication Co., Ltd.(vivo蓝影实验室,vivo移动通信有限公司)

AI总结 本文提出SmartPhotoCrafter,通过统一推理生成过程实现自动摄影图像编辑,提升图像质量与美观,无需人工指令。

Comments tech report

详情
AI中文摘要

传统摄影图像编辑需要用户具备足够的审美理解以提供调整图像质量和相机参数的指令。然而,这种范式依赖于显式的审美意图指令,常存在模糊、不完整或非专家用户无法访问的问题。本文提出SmartPhotoCrafter,将图像编辑视为紧密耦合的推理到生成过程。该模型首先通过Image Critic模块进行图像质量理解并识别缺陷,然后通过Photographic Artist模块实现针对性编辑以提升图像吸引力,无需显式的人类指令。采用多阶段训练流程:(i) 基础预训练以建立基本审美理解和编辑能力,(ii) 适应性训练以多编辑监督结合丰富的语义指导,(iii) 协调推理到生成强化学习以共同优化推理和生成。训练过程中,SmartPhotoCrafter强调逼真图像生成,同时支持图像修复和润色任务,保持颜色和色调相关的语义一致性。我们还构建了阶段特定的数据集,逐步构建推理和可控生成,有效跨模块协作,最终实现高质量的摄影增强。实验表明,SmartPhotoCrafter在自动摄影增强任务中优于现有生成模型,实现了逼真结果并表现出更高的色调敏感性以润色指令。项目页面:https://github.com/vivoCameraResearch/SmartPhotoCrafter.

英文摘要

Traditional photographic image editing typically requires users to possess sufficient aesthetic understanding to provide appropriate instructions for adjusting image quality and camera parameters. However, this paradigm relies on explicit human instruction of aesthetic intent, which is often ambiguous, incomplete, or inaccessible to non-expert users. In this work, we propose SmartPhotoCrafter, an automatic photographic image editing method which formulates image editing as a tightly coupled reasoning-to-generation process. The proposed model first performs image quality comprehension and identifies deficiencies by the Image Critic module, and then the Photographic Artist module realizes targeted edits to enhance image appeal, eliminating the need for explicit human instructions. A multi-stage training pipeline is adopted: (i) Foundation pretraining to establish basic aesthetic understanding and editing capabilities, (ii) Adaptation with reasoning-guided multi-edit supervision to incorporate rich semantic guidance, and (iii) Coordinated reasoning-to generation reinforcement learning to jointly optimize reasoning and generation. During training, SmartPhotoCrafter emphasizes photo-realistic image generation, while supporting both image restoration and retouching tasks with consistent adherence to color- and tone-related semantics. We also construct a stage-specific dataset, which progressively builds reasoning and controllable generation, effective cross-module collaboration, and ultimately high-quality photographic enhancement. Experiments demonstrate that SmartPhotoCrafter outperforms existing generative models on the task of automatic photographic enhancement, achieving photo-realistic results while exhibiting higher tonal sensitivity to retouching instructions. Project page: https://github.com/vivoCameraResearch/SmartPhotoCrafter.

URL PDF HTML 收藏
2604.13054 2026-04-16 cs.CL cs.AI cs.CV

Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling

先caption,后VQA:知识密度而非任务格式驱动多模态扩展

Hongjian Zou, Yue Ge, Qi Ding, Yixuan Liao, Xiaoxin Chen

机构 * vivo AI Lab(vivo人工智能实验室) Wuhan University(武汉大学)

AI总结 研究发现,多模态模型扩展的关键瓶颈在于训练数据的知识密度而非任务格式,通过结构化caption增强和跨模态知识注入可提升模型性能,表明当前多模态模型因训练数据知识覆盖不足而难以扩展。

Comments 23 pages, 4 figures, 10 tables. Preprint

详情
AI中文摘要

多模态大语言模型(MLLMs)虽取得快速进展,但其扩展行为仍不够明确和可预测。增加模型规模和任务多样性往往带来边际效益递减。本文认为,多模态扩展的主要瓶颈不是任务格式,而是训练数据中的知识密度。我们证明,任务特定监督如视觉问答(VQA)提供的语义信息有限,VQA信号可通过caption重构实现几乎无损的性能。随后,我们展示通过结构化caption增强和跨模态知识注入可提升多模态和下游基准性能。在受控实验中,性能更强烈地与语义覆盖相关,而非任务多样性。这些发现表明,当前MLLMs难以扩展的主要原因是训练数据缺乏足够的知识覆盖。我们倡导以知识为中心的多模态训练作为可扩展多模态模型的原理基础。

英文摘要

Multimodal large language models (MLLMs) have achieved rapid progress, yet their scaling behavior remains less clearly characterized and often less predictable than that of text-only LLMs. Increasing model size and task diversity often yields diminishing returns. In this work, we argue that the primary bottleneck in multimodal scaling is not task format, but knowledge density in training data. We first show that task-specific supervision such as Visual Question Answering (VQA) contributes little incremental semantic information beyond image captions: VQA signals can be reconstructed from captions with negligible performance loss. We then demonstrate that increasing knowledge density -- through structured caption enrichment and cross-modal knowledge injection -- leads to consistent performance improvements across multimodal and downstream benchmarks. Across controlled experiments, performance correlates more strongly with semantic coverage than with task diversity. These findings suggest that current MLLMs fail to scale primarily because training data lacks sufficient knowledge coverage. We advocate for knowledge-centric multimodal training as a principled foundation for scalable multimodal models.

URL PDF HTML 收藏
2509.06477 2026-04-16 cs.AI

MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents

MAS-Bench:一种统一的快捷键增强混合移动GUI代理基准测试

Pengxiang Zhao, Guangyi Liu, YaoZhen Liang, Weiqing He, Zhengxi Lu, WenHao Wang, Yuehao Huang, Yuxiang Chai, Zhaolu Kang, Yaxuan Guo, Hao Wang, Kexin Zhang, Liang Liu, Yong Liu

机构 * Zhejiang University(浙江大学) vivo AI Lab(vivo AI实验室) Peking University(北京大学)

AI总结 本文提出MAS-Bench,用于评估结合GUI和快捷键的混合移动自动化代理,包含139个任务和9个评估指标,实验显示混合代理在效率上优于纯GUI代理。

详情
AI中文摘要

本文提出MAS-Bench,用于评估结合GUI和快捷键的混合移动自动化代理,包含139个任务和9个评估指标,实验显示混合代理在效率上优于纯GUI代理。

英文摘要

Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, fostering a promising hybrid paradigm for MLLM-based mobile automation. However, systematic evaluation of GUI-shortcut hybrid agents remains largely underexplored. To bridge this gap, we introduce MAS-Bench, a benchmark that pioneers the evaluation of GUI-shortcut hybrid agents with a specific focus on the mobile domain. Beyond merely using predefined shortcuts, MAS-Bench assesses an agent's capability to autonomously generate shortcuts by discovering and creating reusable, low-cost workflows. It features 139 complex tasks across 11 real-world applications, a knowledge base of 88 predefined shortcuts (APIs, deep-links, RPA scripts), and 9 evaluation metrics. Experiments demonstrate that hybrid agents achieve up to 68.3% success rate and 39% greater execution efficiency than GUI-only counterparts. Furthermore, our evaluation framework effectively reveals the quality gap between predefined and agent-generated shortcuts, validating its capability to assess shortcut generation methods. MAS-Bench addresses the lack of systematic benchmarks for GUI-shortcut hybrid mobile agents, providing a foundational platform for future advancements in creating more efficient and robust intelligent agents. Project page: https://pengxiang-zhao.github.io/MAS-Bench.

URL PDF HTML 收藏
2604.10674 2026-04-14 cs.LG cs.AI cs.CL

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Skill-SD:基于多轮LLM代理的技能条件自蒸馏

Hao Wang, Guozhi Wang, Han Xiao, Yufeng Zhou, Yue Pan, Jichao Wang, Ke Xu, Yafei Wen, Xiaohu Ruan, Xiaoxin Chen, Honggang Qi

机构 * Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院) The Chinese University of Hong Kong(香港中文大学) University of Science and Technology of China(中国科学技术大学) University of Chinese Academy of Sciences(中国科学院大学) vivo AI Lab(vivo AI实验室)

AI总结 Skill-SD通过将代理自身轨迹转化为动态训练监督,结合重要加权反KL损失稳定训练,提升多轮LLM代理性能,实验显示优于标准RL和OPD方法。

Comments Project page: https://k1xe.github.io/skill-sd/

详情
AI中文摘要

强化学习(RL)广泛用于训练多轮交互任务的LLM代理,但其样本效率受限于稀疏奖励和长 horizon。在线自蒸馏(OPSD)通过提供来自拥有真实答案的特权教师的密集标记级监督缓解了这一问题。然而,固定特权信息无法捕捉代理任务中的多样化有效策略,且简单结合OPSD与RL常导致训练崩溃。为解决这些限制,我们引入Skill-SD框架,将代理自身轨迹转化为动态训练监督。完成轨迹被总结为紧凑的自然语言技能,描述成功行为、错误和工作流程。这些技能作为动态特权信息仅条件教师,而学生始终在纯任务提示下行动,并通过蒸馏学习内化指导。为稳定训练,我们推导出重要加权反KL损失以提供梯度正确的标记级蒸馏,并动态同步教师与改进的学生。在代理基准测试中,Skill-SD显著优于标准RL基线,改进了Vanilla GRPO(在AppWorld/Sokoban上分别+14.0%/+10.9%)和Vanilla OPD(+42.1%/+40.6%)。项目页面:https://k1xe.github.io/skill-sd/

英文摘要

Reinforcement learning (RL) has been widely used to train LLM agents for multi-turn interactive tasks, but its sample efficiency is severely limited by sparse rewards and long horizons. On-policy self-distillation (OPSD) alleviates this by providing dense token-level supervision from a privileged teacher that has access to ground-truth answers. However, such fixed privileged information cannot capture the diverse valid strategies in agent tasks, and naively combining OPSD with RL often leads to training collapse. To address these limitations, we introduce Skill-SD, a framework that turns the agent's own trajectories into dynamic training-only supervision. Completed trajectories are summarized into compact natural language skills that describe successful behaviors, mistakes, and workflows. These skills serve as dynamic privileged information conditioning only the teacher, while the student always acts under the plain task prompt and learns to internalize the guidance through distillation. To stabilize the training, we derive an importance-weighted reverse-KL loss to provide gradient-correct token-level distillation, and dynamically synchronize the teacher with the improving student. Experimental results on agentic benchmarks demonstrate that Skill-SD substantially outperforms the standard RL baseline, improving both vanilla GRPO (+14.0%/+10.9% on AppWorld/Sokoban) and vanilla OPD (+42.1%/+40.6%). Project page: https://k1xe.github.io/skill-sd/

URL PDF HTML 收藏
2411.17163 2026-04-14 cs.CV

OSDFace: One-Step Diffusion Model for Face Restoration

OSDFace:面向面部修复的一步扩散模型

Jingkai Wang, Jue Gong, Lin Zhang, Zheng Chen, Xing Liu, Hong Gu, Yutong Liu, Yulun Zhang, Xiaokang Yang

机构 * Shanghai Jiao Tong University(上海交通大学) vivo Mobile Communication Co., Ltd(维沃移动通信有限公司)

AI总结 OSDFace提出一种一步扩散模型,通过视觉嵌入器和面部身份损失提升面部修复的视觉质量和身份一致性。

Comments Accepted to CVPR 2025. The code and model will be available at https://github.com/jkwang28/OSDFace

详情
AI中文摘要

扩散模型在面部修复中表现出色,但多步推理过程计算开销大,限制了实际应用。本文提出OSDFace,一种新的一步扩散模型,通过视觉嵌入器(VRE)更好地捕捉先验信息,并结合面部识别衍生的身份损失确保身份一致性。此外,采用生成对抗网络(GAN)作为指导模型以促进修复面部与真实地面 truth 的分布对齐。实验结果表明,OSDFace在视觉质量和定量指标上均优于现有最先进方法,生成高质量、自然且身份一致的面部图像。代码和模型将在https://github.com/jkwang28/OSDFace发布。

英文摘要

Diffusion models have demonstrated impressive performance in face restoration. Yet, their multi-step inference process remains computationally intensive, limiting their applicability in real-world scenarios. Moreover, existing methods often struggle to generate face images that are harmonious, realistic, and consistent with the subject's identity. In this work, we propose OSDFace, a novel one-step diffusion model for face restoration. Specifically, we propose a visual representation embedder (VRE) to better capture prior information and understand the input face. In VRE, low-quality faces are processed by a visual tokenizer and subsequently embedded with a vector-quantized dictionary to generate visual prompts. Additionally, we incorporate a facial identity loss derived from face recognition to further ensure identity consistency. We further employ a generative adversarial network (GAN) as a guidance model to encourage distribution alignment between the restored face and the ground truth. Experimental results demonstrate that OSDFace surpasses current state-of-the-art (SOTA) methods in both visual quality and quantitative metrics, generating high-fidelity, natural face images with high identity consistency. The code and model will be released at https://github.com/jkwang28/OSDFace.

URL PDF HTML 收藏
2604.08048 2026-04-10 cs.CV

Guiding a Diffusion Model by Swapping Its Tokens

通过交换令牌引导扩散模型

Weijia Zhang, Yuehao Liu, Shanyan Guan, Wu Ran, Yanhao Ge, Wei Li, Chao Ma

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(上海交通大学人工智能研究院教育部人工智能重点实验室) vivo Mobile Communication Co., Ltd.(维沃移动通信有限公司)

AI总结 本文提出一种简单方法,通过交换令牌实现条件和无条件生成的CFG引导,提升图像质量和提示对齐。

Comments Accepted by CVPR 2026 (Oral)

详情
AI中文摘要

分类器自由引导(CFG)是一种广泛使用的推理技术,用于提升扩散模型的图像质量。然而,其依赖文本条件限制了无条件生成的应用。我们提出了一种简单的方法,使CFG引导适用于条件和无条件生成。核心思想是通过简单的令牌交换操作生成扰动预测,并利用其与干净预测之间的方向,引导采样向更高保真度的分布发展。实践中,我们交换空间或通道维度中语义差异最大的令牌潜在表示。与现有方法不同,我们的方法选择性地交换和重组令牌潜在表示,允许更精细地控制扰动及其对生成样本的影响。在MS-COCO 2014、MS-COCO 2017和ImageNet数据集上的实验表明,所提出的自交换引导(SSG)在不同设置下优于先前的无条件方法,在图像保真度和提示对齐方面表现更佳。其细粒度扰动粒度也提高了鲁棒性,减少了扰动强度范围内的副作用。总体而言,SSG将CFG扩展到更广泛的应用,包括条件和无条件生成,并可作为插件轻松插入任何扩散模型以获得即时改进。

英文摘要

Classifier-Free Guidance (CFG) is a widely used inference-time technique to boost the image quality of diffusion models. Yet, its reliance on text conditions prevents its use in unconditional generation. We propose a simple method to enable CFG-like guidance for both conditional and unconditional generation. The key idea is to generate a perturbed prediction via simple token swap operations, and use the direction between it and the clean prediction to steer sampling towards higher-fidelity distributions. In practice, we swap pairs of most semantically dissimilar token latents in either spatial or channel dimensions. Unlike existing methods that apply perturbation in a global or less constrained manner, our approach selectively exchanges and recomposes token latents, allowing finer control over perturbation and its influence on generated samples. Experiments on MS-COCO 2014, MS-COCO 2017, and ImageNet datasets demonstrate that the proposed Self-Swap Guidance (SSG), when applied to popular diffusion models, outperforms previous condition-free methods in image fidelity and prompt alignment under different set-ups. Its fine-grained perturbation granularity also improves robustness, reducing side-effects across a wider range of perturbation strengths. Overall, SSG extends CFG to a broader scope of applications including both conditional and unconditional generation, and can be readily inserted into any diffusion model as a plug-in to gain immediate improvements.

URL PDF HTML 收藏
2604.07955 2026-04-10 cs.LG

Rethinking Residual Errors in Compensation-based LLM Quantization

重新思考基于补偿的LLM量化中的残差误差

Shuaiting Li, Juncan Deng, Kedong Xu, Rongtao Deng, Hong Gu, Minghan Jiang, Haibin Shen, Kejie Huang

机构 * Zhejiang University(浙江大学) vivo Mobile Communication Co., Ltd(维沃移动通信有限公司)

AI总结 本文重新审视残差误差的定义,指出现有方法的校准目标存在次优问题,并引入补偿感知误差以提升LLM量化性能。

Comments ICLR'26 camera ready

详情
AI中文摘要

基于权重补偿的方法通过迭代应用量化和权重补偿来最小化输出误差,最近在量化大型语言模型(LLMs)中表现出色。GPTQ引入了关键技术使此类迭代方法在具有数十亿参数的LLMs中成为现实。GPTAQ通过引入不对称校准过程对齐每个量化层的输出与其全精度对应物,并将残差误差纳入权重补偿框架。本文重新审视残差误差的公式。我们发现现有方法存在次优校准目标:在层内校准过程中,它们将量化输出与补偿权重的输出对齐,而非原始全精度模型的真实输出。因此,我们重新定义目标,使量化模型的输出与全精度模型的原始输出在每一步精确对齐。然后揭示残差误差不仅来自前一层的输出差异,还来自各层内补偿权重与原始权重之间的差异,我们将其命名为'补偿感知误差'。通过继承GPTAQ中的神经分解技术,我们能够高效地将此补偿感知误差纳入权重更新过程中。在各种LLM和量化设置上的广泛实验表明,所提出的改进能够无缝集成到GPTQ和GPTAQ中,显著提升其量化性能。我们的代码在https://github.com/list0830/ResComp上公开可用。

英文摘要

Methods based on weight compensation, which iteratively apply quantization and weight compensation to minimize the output error, have recently demonstrated remarkable success in quantizing Large Language Models (LLMs). The representative work, GPTQ, introduces several key techniques that make such iterative methods practical for LLMs with billions of parameters. GPTAQ extends this approach by introducing an asymmetric calibration process that aligns the output of each quantized layer with its full-precision counterpart, incorporating a residual error into the weight compensation framework. In this work, we revisit the formulation of the residual error. We identify a sub-optimal calibration objective in existing methods: during the intra-layer calibration process, they align the quantized output with the output from compensated weights, rather than the true output from the original full-precision model. Therefore, we redefine the objective to precisely align the quantized model's output with the original output of the full-precision model at each step. We then reveal that the residual error originates not only from the output difference of the preceding layer but also from the discrepancy between the compensated and original weights within each layer, which we name the 'compensation-aware error'. By inheriting the neuron decomposition technique from GPTAQ, we can efficiently incorporate this compensation-aware error into the weight update process. Extensive experiments on various LLMs and quantization settings demonstrate that our proposed enhancements integrate seamlessly with both GPTQ and GPTAQ, significantly improving their quantization performance. Our code is publicly available at https://github.com/list0830/ResComp.

URL PDF HTML 收藏
2604.07363 2026-04-10 cs.LG

Benchmark Shadows: Data Alignment, Parameter Footprints, and Generalization in Large Language Models

基准阴影:大型语言模型中的数据对齐、参数足迹与泛化

Hongjian Zou, Yidan Wang, Qi Ding, Yixuan Liao, Xiaoxin Chen

机构 * Vivo AI Lab, Shenzhen, China(维沃人工智能实验室,深圳,中国) Hong Kong University of Science and Technology, Hong Kong, China(香港科技大学,香港,中国)

AI总结 本文研究了大型语言模型中基准表现与广度能力之间的差异,通过数据干预发现数据分布影响参数适应与泛化能力,提出参数空间诊断方法揭示不同训练模式的结构特征。

Comments 28 pages, 26 figures, 8 tables

详情
AI中文摘要

大型语言模型常常在基准测试中取得显著提升,但并未在更广泛的能力上同步进步。我们假设这种差异源于数据分布带来的训练制度差异。为此,我们设计了受控的数据干预,以固定训练设置下隔离分布效应。研究发现,与基准对齐的数据提高了狭窄评估指标,但限制了更广泛表征的发展;而扩展覆盖的数据导致更分布化的参数适应和更好的泛化。我们进一步引入基于谱分析和秩分析的参数空间诊断,揭示了这些模式的结构特征。在多样化的开源模型家族中观察到相似模式,包括多模态模型作为关键案例研究,表明这些效应超出了受控环境。对提示重复的案例研究表明,并非所有数据伪影都会引发模式转变。这些结果表明,仅凭基准表现不足以描述模型能力,并突显了数据分布在塑造学习动态中的重要性。

英文摘要

Large language models often achieve strong benchmark gains without corresponding improvements in broader capability. We hypothesize that this discrepancy arises from differences in training regimes induced by data distribution. To investigate this, we design controlled data interventions that isolate distributional effects under fixed training settings. We find that benchmark-aligned data improves narrow evaluation metrics while limiting broader representational development, whereas coverage-expanding data leads to more distributed parameter adaptation and better generalization. We further introduce parameter-space diagnostics based on spectral and rank analyses, which reveal distinct structural signatures of these regimes. Similar patterns are observed across diverse open-source model families, including multimodal models as a key case study, suggesting that these effects extend beyond controlled settings. A case study on prompt repetition shows that not all data artifacts induce regime shifts. These results indicate that benchmark performance alone is insufficient to characterize model capability, and highlight the importance of data distribution in shaping learning dynamics.

URL PDF HTML 收藏
2602.01554 2026-04-07 cs.LG cs.AI cs.CV

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs

InfoTok: 基于信息理论的容量受限共享视觉标记化在统一多模态大语言模型中的应用

Lv Tang, Tianyi Zheng, Bo Li, Xingyu Li

机构 * University of Alberta(阿尔伯塔大学) vivo Mobile Communication Co., Ltd(维沃移动通信有限公司)

AI总结 本文提出InfoTok,一种基于信息瓶颈原理的信息正则化标记化机制,通过互信息约束平衡压缩与任务相关性,提升统一多模态大语言模型的图像理解和生成性能。

详情
AI中文摘要

统一多模态大语言模型(MLLMs)旨在在一个框架内统一图像理解和生成,其中共享视觉标记器作为唯一接口,将高维图像映射到有限的标记预算中,以支持下游多模态推理和生成。然而,现有共享标记设计大多受架构驱动,缺乏明确标准来决定应保留哪些信息以同时支持语义抽象和视觉细节。本文采用容量受限视角,将共享标记器视为计算受限的学习者,其有限的表示预算应优先考虑可重用的结构而非难以利用的高熵变化和冗余。受此观点启发,我们提出InfoTok,一种基于信息瓶颈(IB)原理的信息正则化标记化机制。InfoTok通过施加互信息(MI)约束,控制从图像到共享标记到多模态输出的信息流,从而在压缩与任务相关性之间实现有原则的权衡,同时鼓励跨模态一致性。由于MI对高维视觉表示不可行,我们用实际可微的依赖估计器实例化InfoTok,包括变分IB公式和基于希尔伯特-施密特独立准则(HSIC)的替代方案。将InfoTok集成到三个代表性的统一MLLMs中,无需引入额外训练数据,一致提升了图像理解和生成性能。这些结果支持信息正则化的视觉标记化作为统一MLLMs中标记学习的坚实基础。

英文摘要

Unified multimodal large language models (MLLMs) aim to unify image understanding and image generation within a single framework, where a shared visual tokenizer serves as the sole interface that maps high-dimensional images into a limited token budget for downstream multimodal reasoning and synthesis. However, existing shared-token designs are largely architecture-driven and lack an explicit criterion for what information should be preserved to simultaneously support semantic abstraction and visual detail. In this paper, we adopt a capacity-constrained perspective, viewing the shared tokenizer as a compute-bounded learner whose finite representational budget should prioritize reusable structure over hard-to-exploit high-entropy variations and redundancy. Motivated by this view, we propose \textbf{\textit{InfoTok}}, an information-regularized tokenization mechanism grounded in the Information Bottleneck (IB) principle. InfoTok explicitly controls information flow from images to shared tokens to multimodal outputs by imposing mutual-information (MI) constraints that enforce a principled trade-off between compression and task relevance, while also encouraging cross-modal consistency. Because MI is intractable for high-dimensional visual representations, we instantiate InfoTok with practical, differentiable dependence estimators, including a variational IB formulation and a Hilbert Schmidt Independence Criterion (HSIC) based alternative. Integrated into three representative unified MLLMs without introducing any additional training data, InfoTok consistently improves both image understanding and generation performance. These results support information-regularized visual tokenization as a sound basis for token learning in unified MLLMs.

URL PDF HTML 收藏
2511.18886 2026-03-19 cs.CV

MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration

MagicWorld: 向交互视频世界探索的长时稳定迈进

Guangyuan Li, Bo Li, Jinwei Chen, Xiaobin Hu, Lei Zhao, Peng-Tao Jiang

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) vivo BlueImage Lab, vivo Mobile Communication Co., Ltd.(vivo蓝影实验室,vivo移动通信有限公司) National University of Singapore(新加坡国立大学)

AI总结 本文提出MagicWorld模型,通过引入流引导运动保留约束和增强交互训练策略,解决复杂环境中运动漂移和长时交互误差累积问题,并构建RealWM120K数据集验证其有效性。

详情
AI中文摘要

本文提出MagicWorld模型,通过引入流引导运动保留约束和增强交互训练策略,解决复杂环境中运动漂移和长时交互误差累积问题,并构建RealWM120K数据集验证其有效性。

英文摘要

Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First, they exhibit motion drift in complex environments with multiple interacting subjects, where dynamic subjects fail to follow realistic motion patterns during scene evolution. Second, they suffer from error accumulation in long-horizon interactions, where autoregressive generation gradually drifts from earlier scene states and causes structural and semantic inconsistencies. In this paper, we propose MagicWorld, an interactive video world model built upon an autoregressive framework. To address motion drift, we incorporate a flow-guided motion preservation constraint that mitigates motion degradation in dynamic subjects, encouraging realistic motion patterns and stable interactions during scene evolution. To mitigate error accumulation in long-horizon interactions, we design two complementary strategies, including a history cache retrieval strategy and an enhanced interactive training strategy. The former reinforces historical scene states by retrieving past generations during interaction, while the latter adopts multi-shot aggregated distillation with dual-reward weighting for interactive training, enhancing long-term stability and reducing error accumulation. In addition, we construct RealWM120K, a real-world dataset with diverse city-walk videos and multimodal annotations to support dynamic perception and long-horizon world modeling. Experimental results demonstrate that MagicWorld improves motion realism and alleviates error accumulation during long-horizon interactions.

URL PDF HTML 收藏
2505.22977 2026-03-19 cs.CV

HyperMotionX: The Dataset and Benchmark with DiT-Based Pose-Guided Human Image Animation of Complex Motions

HyperMotionX:基于DiT的Pose引导人体图像动画的Dataset和Benchmark

Shuolin Xu, Siming Zheng, Ziyi Wang, HC Yu, Jinwei Chen, Huaqi Zhang, Daquan Zhou, Tong-Yee Lee, Bo Li, Peng-Tao Jiang

机构 * Bournemouth University(伯恩茅斯大学) vivo BlueImage Lab, vivo Mobile Communication Co., Ltd(vivo 蓝影实验室,vivo 通信有限公司) Peking University(北京大学) National Cheng-Kung University(国立成功大学)

AI总结 本文提出基于DiT的高效人体动画生成基线和改进的RoPE模块,构建了HyperMotionX数据集和基准,提升了复杂人体运动动画的结构稳定性和外观一致性。

Comments 17 pages, 7 figures

详情
AI中文摘要

最近扩散模型的进步显著提高了条件视频生成,特别是在姿态引导的人体图像动画任务中。尽管现有方法能够生成高保真且时间一致的动画序列,在常规运动和静态场景中表现良好,但面对包含高度动态、非标准运动的复杂人体运动时仍存在明显局限,缺乏高质量的评估基准。为解决这一挑战,我们提出了一种简洁而强大的基于DiT的人体动画生成基线,并设计了空间低频增强RoPE模块,通过引入可学习的频率缩放来选择性增强低频空间特征建模。此外,我们引入了Open-HyperMotionX数据集和HyperMotionX Bench,提供了高质量的人体姿态标注和精选视频片段,用于在复杂人体运动条件下评估和改进姿态引导的人体图像动画模型。我们的方法显著提高了高度动态人体运动序列的结构稳定性和外观一致性。大量实验验证了我们数据集和所提方法在提升复杂人体运动图像动画生成质量方面的有效性。代码、模型权重和数据集已公开在https://vivocameraresearch.github.io/hypermotion/

英文摘要

Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods are capable of generating high-fidelity and time-consistent animation sequences in regular motions and static scenes. However there are still obvious limitations when facing complex human body motions that contain highly dynamic, non-standard motions, and the lack of a high-quality benchmark for evaluation of complex human motion animations. To address this challenge, we propose a concise yet powerful DiT-based human animation generation baseline and design spatial low-frequency enhanced RoPE, a novel module that selectively enhances low-frequency spatial feature modeling by introducing learnable frequency scaling. Furthermore, we introduce the Open-HyperMotionX Dataset and HyperMotionX Bench, which provide high-quality human pose annotations and curated video clips for evaluating and improving pose-guided human image animation models under complex human motion conditions. Our method significantly improves structural stability and appearance consistency in highly dynamic human motion sequences. Extensive experiments demonstrate the effectiveness of our dataset and proposed approach in advancing the generation quality of complex human motion image animations. The codes, model weights, and dataset have been made publicly available at https://vivocameraresearch.github.io/hypermotion/

URL PDF HTML 收藏
2603.14916 2026-03-17 cs.CV cs.MM

EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing

EditHF-1M: 一个百万级的丰富人类偏好反馈用于图像编辑

Zitong Xu, Huiyu Duan, Zhongpeng Ji, Xinyun Zhang, Yutao Liu, Xiongkuo Min, Ke Gu, Jian Zhang, Shusong Xu, Jinwei Chen, Bo Li, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Vivo Mobile Communication Co., Ltd(vivo移动通信有限公司) University of Electronic and Science Technology of China(电子科技大学) Ocean University of China(海洋大学) Beijing University of Technology(北京理工大学)

AI总结 本文提出EditHF-1M百万级图像编辑数据集及基于其的EditHF评估模型和EditHF-Reward奖励模型,通过强化学习优化文本引导图像编辑模型,实验表明其在对齐人类偏好和泛化能力方面表现优异。

详情
AI中文摘要

近期文本引导图像编辑(TIE)模型取得了显著进展,但许多编辑图像仍存在伪影、意外编辑和不美观内容等问题。尽管已提出一些评估编辑图像的基准和方法,但可扩展的评估模型仍然缺乏,限制了图像编辑的人类反馈奖励模型的发展。为解决这些挑战,我们首先引入EditHF-1M,一个包含超过2900万对人类偏好和148000个人类平均意见评分的百万级图像编辑数据集,两者均从三个维度评估:视觉质量、指令对齐性和属性保留。基于EditHF-1M,我们提出EditHF,一个基于多模态大语言模型(MLLM)的评估模型,以提供图像编辑的人类对齐反馈。最后,我们引入EditHF-Reward,利用EditHF作为奖励信号,通过强化学习优化文本引导图像编辑模型。大量实验表明,EditHF在对齐人类偏好和在其他数据集上的泛化能力方面表现优异。此外,我们使用EditHF-Reward微调Qwen-Image-Edit,取得了显著的性能提升,这证明了EditHF作为奖励模型的能力,能够扩展图像编辑。数据集和代码将在我们的GitHub仓库中发布:https://github.com/IntMeGroup/EditHF。

英文摘要

Recent text-guided image editing (TIE) models have achieved remarkable progress, while many edited images still suffer from issues such as artifacts, unexpected editings, unaesthetic contents. Although some benchmarks and methods have been proposed for evaluating edited images, scalable evaluation models are still lacking, which limits the development of human feedback reward models for image editing. To address the challenges, we first introduce \textbf{EditHF-1M}, a million-scale image editing dataset with over 29M human preference pairs and 148K human mean opinion ratings, both evaluated from three dimensions, \textit{i.e.}, visual quality, instruction alignment, and attribute preservation. Based on EditHF-1M, we propose \textbf{EditHF}, a multimodal large language model (MLLM) based evaluation model, to provide human-aligned feedback from image editing. Finally, we introduce \textbf{EditHF-Reward}, which utilizes EditHF as the reward signal to optimize the text-guided image editing models through reinforcement learning. Extensive experiments show that EditHF achieves superior alignment with human preferences and demonstrates strong generalization on other datasets. Furthermore, we fine-tune the Qwen-Image-Edit using EditHF-Reward, achieving significant performance improvements, which demonstrates the ability of EditHF to serve as a reward model to scale-up the image editing. Both the dataset and code will be released in our GitHub repository: https://github.com/IntMeGroup/EditHF.

URL PDF HTML 收藏
2509.05554 2026-03-09 cs.CV cs.IR

RED: Robust Event-Guided Motion Deblurring with Modality-Specific Disentanglement

RED: 基于事件引导的鲁棒运动去模糊化与模态特定解耦

Yihong Leng, Siming Zheng, Jinwei Chen, Bo Li, Jiaojiao Li, Peng-Tao Jiang

机构 * Xidian University(西安电子科技大学) vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司)

AI总结 RED通过事件引导的鲁棒运动去模糊化方法,结合模态特定解耦与选择性融合,提升模糊图像的清晰度和鲁棒性。

详情
AI中文摘要

事件引导的运动去模糊化利用事件相机的高时间分辨率运动线索来重建清晰图像。然而,在实际捕获中,阈值引起的事件欠报会导致运动线索缺失和碎片化,现有方法在此情况下性能下降,主要受限于两个方面:一是对密集且稳定的事件的假设,二是模态不加区分的提取和融合,无法将有用的运动线索与受损事件分开,导致它们污染跨模态表示。在本文中,我们首先引入了一种面向鲁棒性的扰动策略(RPS),模拟动态视觉传感器的各种触发阈值,使我们的模型暴露于多样的欠报模式中,从而在未知条件下提高鲁棒性。在此基础上,我们提出RED,一个鲁棒的事件引导去模糊化网络,遵循“先解耦再选择性融合”的原则。具体来说,模态特定的表示机制将输入解耦为图像语义、事件运动和跨模态表示,分别捕捉外观、运动和互补交互。利用可靠的解耦特征,我们选择性地融合模态以增强模糊图像中对运动敏感的区域,并用语义上下文丰富欠报事件。在合成和真实世界数据集上的广泛实验表明,RED在准确性和鲁棒性方面均实现了最先进的性能。

英文摘要

Event-guided motion deblurring reconstructs sharp images using the high-temporal-resolution motion cues from event cameras. However, in real capture, thresholding-induced event under-reporting causes missing and fragmented motion cues, under which existing methods often degrade in performance due to two limitations: i) assumptions of dense and stable events, and ii) modality-indiscriminate extraction and fusion that fail to separate useful motion cues from disrupted events, allowing them to contaminate cross-modal representations. In this paper, we first introduce a Robustness-Oriented Perturbation Strategy (RPS) that mimics various trigger thresholds of dynamic vision sensors, exposing our model to diverse under-reporting patterns and thereby improving robustness under unknown conditions. Built upon this setting, we propose RED, a Robust Event-guided Deblurring network, following the principle of disentangle first and then fuse selectively. Specifically, the Modality-specific Representation Mechanism disentangles the inputs into image-semantic, event-motion, and cross-modal representations, capturing appearance, motion, and complementary interactions, respectively. With the reliable disentangled features, we selectively fuse modalities to enhance motion-sensitive areas in blurry images and enrich under-reported events with semantic context. Extensive experiments on synthetic and real-world datasets demonstrate RED consistently achieves state-of-the-art performance in terms of both accuracy and robustness.

URL PDF HTML 收藏
2410.09864 2026-03-09 cs.CV

AuthFace: Towards Authentic Blind Face Restoration with Face-oriented Generative Diffusion Prior

AuthFace: 向真实盲脸修复迈进的面向面部生成扩散先验

Guoqiang Liang, Qingnan Fan, Bingtao Fu, Jinwei Chen, Hong Gu, Lin Wang

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) vivo Mobile Communication Co., Ltd(vivo移动通信有限公司) Nanyang Technological University(南洋理工大学)

AI总结 AuthFace通过面向面部的生成扩散先验和摄影引导标注,实现高质量盲脸修复,提升面部细节还原和减少关键区域伪影。

Comments ACM MM 25, Codes and datasets are available at https://github.com/EthanLiang99/AuthFace

详情
AI中文摘要

盲脸修复(BFR)是计算机视觉中的基本且具有挑战性的问题。为了从低质量图像中忠实恢复高质量(HQ)照片,最近的研究主要依赖于强大预训练文本到图像(T2I)扩散模型的面部图像先验。然而,此类先验往往导致非面部特征的错误生成和面部细节不足,因此在实际应用中不够实用。在本文中,我们提出了一种新的框架,即AuthFace,通过探索面向面部的生成扩散先验来实现高度真实的面部修复结果。为了学习此类先验,我们首先收集了1500张高质量图像,分辨率超过8K,由专业摄影师拍摄。基于该数据集,我们引入了一种新的面向面部的修复微调流程,对预训练的T2I模型进行微调。识别质量优先和摄影引导的标注关键标准,我们邀请摄影师对展示丰富面部特征的高质量图像进行润色和审核。摄影引导的标注系统充分挖掘了这些高质量摄影图像的潜力。通过这种方式,预训练T2I扩散模型的自然图像先验可以被微妙地利用,特别是增强其在面部细节修复方面的能力。此外,为了减少关键面部区域(如眼睛和嘴巴)中的伪影,我们提出了一种时间感知的潜在面部特征损失,以学习真实的面部修复过程。在合成和真实世界BFR数据集上的广泛实验表明了我们方法的优越性。

英文摘要

Blind face restoration (BFR) is a fundamental and challenging problem in computer vision. To faithfully restore high-quality (HQ) photos from poor-quality ones, recent research endeavors predominantly rely on facial image priors from the powerful pretrained text-to-image (T2I) diffusion models. However, such priors often lead to the incorrect generation of non-facial features and insufficient facial details, thus rendering them less practical for real-world applications. In this paper, we propose a novel framework, namely AuthFace that achieves highly authentic face restoration results by exploring a face-oriented generative diffusion prior. To learn such a prior, we first collect a dataset of 1.5K high-quality images, with resolutions exceeding 8K, captured by professional photographers. Based on the dataset, we then introduce a novel face-oriented restoration-tuning pipeline that fine-tunes a pretrained T2I model. Identifying key criteria of quality-first and photography-guided annotation, we involve the retouching and reviewing process under the guidance of photographers for high-quality images that show rich facial features. The photography-guided annotation system fully explores the potential of these high-quality photographic images. In this way, the potent natural image priors from pretrained T2I diffusion models can be subtly harnessed, specifically enhancing their capability in facial detail restoration. Moreover, to minimize artifacts in critical facial areas, such as eyes and mouth, we propose a time-aware latent facial feature loss to learn the authentic face restoration process. Extensive experiments on the synthetic and real-world BFR datasets demonstrate the superiority of our approach.

URL PDF HTML 收藏
2511.19524 2026-03-05 cs.CV cs.MA

VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning

VideoChat-M1: 通过多智能体强化学习实现视频理解的协作策略规划

Boyu Chen, Zikang Wang, Zhengrong Yue, Kainan Yan, Chenyun Yu, Yi Huang, Zijun Liu, Yafei Wen, Xiaoxin Chen, Yang Liu, Peng Li, Yali Wang

机构 * Shenzhen Key Lab of Computer Vision and Pattern Recognition(深圳计算机视觉与模式识别重点实验室) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) VIVO AI Lab(VIVO人工智能实验室) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shenzhen Campus of Sun Yat-sen University(孙逸仙大学深圳校区) Shanghai Jiao Tong University(上海交通大学) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University(清华大学计算机科学与技术系,人工智能研究院)

AI总结 VideoChat-M1通过多智能体强化学习实现视频理解的协作策略规划,显著提升复杂视频任务的性能。

Comments Accepted by CVPR 2026

详情
AI中文摘要

通过利用工具增强的多模态大语言模型(MLLMs),多智能体框架正在推动视频理解的进步。然而,大多数方法采用静态且不可学习的工具调用机制,这限制了发现对于时间或空间复杂视频中必要的多样化线索。为了解决这一挑战,我们提出了一种新的视频理解多智能体系统,即VideoChat-M1。与使用单一或固定策略不同,VideoChat-M1采用了一种独特的协作策略规划(CPP)范式,包含三个关键过程。(1)策略生成:每个智能体生成其独特的工具调用策略,以适应用户的查询;(2)策略执行:每个智能体依次调用相关工具以执行其策略并探索视频内容;(3)策略通信:在策略执行的中间阶段,智能体相互交互以更新各自的策略。通过这种协作框架,所有智能体协同工作,根据同伴提供的上下文洞察动态改进各自的首选策略,以有效响应用户的查询。此外,我们为CPP范式配备了简洁的多智能体强化学习(MARL)方法。因此,策略智能体团队可以联合优化以提升VideoChat-M1的性能,由最终答案奖励和中间协作过程反馈引导。大量实验表明,VideoChat-M1在八个基准测试中四个任务上均取得了SOTA性能。值得注意的是,在LongVideoBench上,我们的方法比SOTA模型Gemini 2.5 Pro高出3.6%,比GPT-4o高出15.6%。

英文摘要

By leveraging tool-augmented Multimodal Large Language Models (MLLMs), multi-agent frameworks are driving progress in video understanding. However, most of them adopt static and non-learnable tool invocation mechanisms, which limit the discovery of diverse clues essential for robust perception and reasoning regarding temporally or spatially complex videos. To address this challenge, we propose a novel Multi-agent system for video understanding, namely VideoChat-M1. Instead of using a single or fixed policy, VideoChat-M1 adopts a distinct Collaborative Policy Planning (CPP) paradigm with multiple policy agents, which comprises three key processes. (1) Policy Generation: Each agent generates its unique tool invocation policy tailored to the user's query; (2) Policy Execution: Each agent sequentially invokes relevant tools to execute its policy and explore the video content; (3) Policy Communication: During the intermediate stages of policy execution, agents interact with one another to update their respective policies. Through this collaborative framework, all agents work in tandem, dynamically refining their preferred policies based on contextual insights from peers to effectively respond to the user's query. Moreover, we equip our CPP paradigm with a concise Multi-Agent Reinforcement Learning (MARL) method. Consequently, the team of policy agents can be jointly optimized to enhance VideoChat-M1's performance, guided by both the final answer reward and intermediate collaborative process feedback. Extensive experiments demonstrate that VideoChat-M1 achieves SOTA performance across eight benchmarks spanning four tasks. Notably, on LongVideoBench, our method outperforms the SOTA model Gemini 2.5 pro by 3.6% and GPT-4o by 15.6%.

URL PDF HTML 收藏