arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-12-15 至 2025-12-15 共收录 61 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3 篇

2512.11749 2025-12-15 cs.CV 89%

SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder

SVG-T2I: 在视觉基础模型表示空间中扩展文本到图像的潜在扩散模型而不使用变分自编码器

Minglei Shi, Haolin Wang, Borui Zhang, Wenzhao Zheng, Bohan Zeng, Ziyang Yuan, Xiaoshi Wu, Yuanxing Zhang, Huan Yang, Xintao Wang, Pengfei Wan, Kun Gai, Jie Zhou, Jiwen Lu

机构 * Department of Automation, Tsinghua University(自动化系,清华大学) Kling Team, Kuaishou Technology(快手技术团队)

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 SVG-T2I通过在视觉基础模型表示空间中直接进行文本到图像生成,实现了高质量的图像合成并验证了VFM在生成任务中的能力。

Comments Code Repository: https://github.com/KlingTeam/SVG-T2I; Model Weights: https://huggingface.co/KlingTeam/SVG-T2I

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16713 2025-12-15 cs.CV 89%

Conditional Text-to-Image Generation with Reference Guidance

基于参考引导的条件文本到图像生成

Taewook Kim, Ze Wang, Zhengyuan Yang, Jiang Wang, Lijuan Wang, Zicheng Liu, Qiang Qiu

机构 * Purdue University(普渡大学) AMD Microsoft(微软)

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 本文提出基于参考引导的条件文本到图像生成方法,通过专家插件提升模型在文本拼写和多语言生成等任务上的表现。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14275 2025-12-15 cs.CV 85%

Free-Lunch Color-Texture Disentanglement for Stylized Image Generation

无需调优的色彩-纹理解耦用于风格化图像生成

Jiang Qin, Senmao Li, Alexandra Gomez-Villa, Shiqi Yang, Yaxing Wang, Kai Wang, Joost van de Weijer

机构 * Harbin Institute of Technology, China(哈尔滨工业大学) VCIP, CS, Nankai University, China(南开大学) Computer Vision Center, Spain(西班牙计算机视觉中心) Universitat Autònoma de Barcelona, Spain(巴塞罗那自治大学) Program of Computer Science, City University of Hong Kong (Dongguan), China(香港城市大学(东莞)计算机科学系) City University of Hong Kong, HK SAR, China(香港城市大学)

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV

AI总结 本文提出无需调优的SADis方法,通过分离颜色-纹理嵌入提升风格化图像生成的精度与定制能力。

Comments Accepted by NeurIPS2025. Code is available at https://deepffff.github.io/sadis.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 图像编辑 2 篇

2512.11395 2025-12-15 cs.CV 83%

FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing

FlowDC: 基于流的解耦-衰减用于复杂图像编辑

Yilei Jiang, Zhen Wang, Yanghao Wang, Jun Yu, Yueting Zhuang, Jun Xiao, Long Chen

机构 * Zhejiang University(浙江大学) HKUST(香港科技大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 图像编辑 :image editing(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 FlowDC通过解耦和衰减技术提升复杂图像编辑的源一致性与语义对齐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00269 2025-12-15 cs.GR cs.CV 62%

3D-LATTE: Latent Space 3D Editing from Textual Instructions

3D-LATTE: 从文本指令进行3D编辑的潜在空间方法

Maria Parelli, Michael Oechsle, Michael Niemeyer, Federico Tombari, Andreas Geiger

机构 * University of Tübingen, Tübingen AI Center(图宾根大学,图宾根人工智能中心) Google Zurich(谷歌瑞士总部)

专题命中 图像编辑 :diffusion(abstract);分类 cs.CV、cs.GR

AI总结 3D-LATTE通过在潜在空间中直接操控3D几何,实现高质量的文本指令驱动3D编辑。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 扩散模型 47 篇

2512.11464 2025-12-15 cs.CV cs.AI cs.LG 85%

Exploring MLLM-Diffusion Information Transfer with MetaCanvas

探索MLLM-扩散信息传输与MetaCanvas

Han Lin, Xichen Pan, Ziqi Huang, Ji Hou, Jialiang Wang, Weifeng Chen, Zecheng He, Felix Juefei-Xu, Junzhe Sun, Zhipeng Fan, Ali Thabet, Mohit Bansal, Chu Wang

机构 * Meta Superintelligence Labs(Meta 超智能实验室) New York University(纽约大学) Nanyang Technological University(南洋理工大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 MetaCanvas通过使MLLMs在潜在空间中推理和规划,提升多模态生成的精确度和结构化控制。

Comments Project page: https://metacanvas.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10041 2025-12-15 cs.CV cs.AI 83%

MetaVoxel: Joint Diffusion Modeling of Imaging and Clinical Metadata

MetaVoxel:联合图像与临床元数据的扩散建模

Yihao Liu, Chenyu Gao, Lianrui Zuo, Michael E. Kim, Brian D. Boyd, Lisa L. Barnes, Walter A. Kukull, Lori L. Beason-Held, Susan M. Resnick, Timothy J. Hohman, Warren D. Taylor, Bennett A. Landman

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 MetaVoxel通过联合扩散建模统一图像与临床元数据,实现图像生成、年龄估计和性别预测,性能媲美传统任务特定模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13387 2025-12-15 cs.CV cs.AI 83%

Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model

通用去噪扩散代码本模型(gDDCM):利用预训练扩散模型进行图像分词

Fei Kong

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 gDDCM通过统一框架和回溯策略提升图像分词效率,实现对主流扩散模型的兼容性与高质量重建

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11763 2025-12-15 cs.CV 79%

Reducing Domain Gap with Diffusion-Based Domain Adaptation for Cell Counting

通过扩散域适应减少领域差距用于细胞计数

Mohammad Dehghanmanshadi, Wallapak Tavanapong

机构 * Computer Science Department , Iowa State University, IA, USA(计算机科学系,爱荷华州立大学,IA,USA)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本研究通过基于InST的风格迁移方法,有效减少合成与真实显微镜数据间的领域差距,提升细胞计数性能,同时减少人工标注工作量。

Comments Accepted at ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11274 2025-12-15 cs.CV 79%

FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion

FilmWeaver: 通过缓存引导自回归扩散编织一致的多镜头视频

Xiangyang Luo, Qingyu Li, Xiaokun Liu, Wenyu Qin, Miao Yang, Meng Wang, Pengfei Wan, Di Zhang, Kun Gai, Shao-Lun Huang

机构 * Kling Team, Kuaishou Technology(快手科技 Kling Team)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FilmWeaver通过缓存引导自回归扩散模型实现多镜头视频的一致性生成,支持灵活交互和多任务应用。

Comments AAAI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07500 2025-12-15 cs.CV 79%

MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer

多主体视频动作迁移:通过视频扩散变换器实现

Penghui Liu, Jiangshan Wang, Yutong Shen, Shanhui Mo, Chenyang Qi, Yue Ma

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 MultiMotion通过视频扩散变换器实现多主体视频动作迁移,引入AMF和RectPC提升动作解纠缠与生成效率,构建首个专用基准数据集。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07410 2025-12-15 cs.CV 79%

InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs

InterAgent:基于物理的多智能体命令执行通过交互图上的扩散

Bin Li, Ruichi Zhang, Han Liang, Jingyan Zhang, Juze Zhang, Xin Chen, Lan Xu, Jingyi Yu, Jingya Wang

机构 * ShanghaiTech University(上海科技大学) University of Pennsylvania(宾夕法尼亚大学) ByteDance(字节跳动) Stanford University(斯坦福大学) InstAdapt

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 InterAgent通过交互图上的扩散模型实现了基于物理的多智能体协调控制,能够从文本提示中生成连贯且物理合理的多代理行为。

Comments Project page: https://binlee26.github.io/InterAgent-Page

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11330 2025-12-15 cs.CV 79%

Noise Matters: Optimizing Matching Noise for Diffusion Classifiers

噪声至关重要:优化扩散分类器的匹配噪声

Yanghao Wang, Long Chen

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出NoOp方法,通过频率匹配和空间匹配原则优化扩散分类器的噪声,以提升分类稳定性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23834 2025-12-15 eess.IV cs.CV 79%

Denoising Diffusion Models for Anomaly Localization in Medical Images

医学图像中异常定位的去噪扩散模型

Cosmin I. Bercea, Philippe C. Cattin, Julia A. Schnabel, Julia Wolleb

机构 * School of Computation, Information and Technology, Technical University of Munich(慕尼黑技术大学计算、信息与技术学院) Institute of Machine Learning in Biomedical Imaging, Helmholtz Munich, Germany(生物医学影像机器学习研究所,海德堡慕尼黑德国) Department of Biomedical Engineering, University of Basel, Allschwil, Switzerland(巴塞尔大学生物医学工程系,瑞士Allschwil) School of Biomedical Engineering and Imaging Sciences, King’s College London, United Kingdom(伦敦国王学院生物医学工程与影像科学学院,英国) Department of Biomedical Informatics and Data Science, Yale University School of Medicine, New Haven, CT, USA(耶鲁大学医学院生物医学信息学与数据科学系,美国新罕布什尔州新 Haven) Yale Biomedical Imaging Institute, Yale University, New Haven, CT, USA(耶鲁大学生物医学影像研究院,美国新罕布什尔州新 Haven)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文探讨了利用去噪扩散模型在医学图像中实现异常定位的方法,分析了不同监督方案的有效性及挑战,指出扩散模型在该领域的应用潜力。

Comments Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://melba-journal.org/2025:030

Journal ref Machine.Learning.for.Biomedical.Imaging. 3 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11721 2025-12-15 math.AP 78%

Stability of stationary reaction diffusion-degenerate Nagumo fronts I: spectral analysis

静止反应扩散-退化Nagumo前沿的稳定性 I:谱分析

Raffaele Folino, César A. Hernández Melo, Luis F. López Ríos, Ramón G. Plaza

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文研究了反应扩散方程中退化Nagumo型前沿的谱稳定性,通过谱分析和能量估计证明了线性化算子的实谱和谱间隙,以及原点的简单孤立特征值。

Comments 35 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11511 2025-12-15 q-bio.QM 78%

DREAM-B3P: Dual-Stream Transformer Network Enhanced by Feedback Diffusion Model for Blood-Brain Barrier Penetrating Peptide Prediction

DREAM-B3P:通过反馈扩散模型增强的双流Transformer网络用于血脑屏障穿透肽预测

Kaijie Wang, Le Yin, Aodi Tian, Zhiqiang Wei, Zai Yang, Min Han, Qichun Wei, Sheng Wang

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 DREAM-B3P通过结合反馈扩散模型和双流Transformer,有效缓解数据不平衡问题,提升血脑屏障穿透肽预测的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11428 2025-12-15 math.OC math.CV 78%

A new $ν$-metric computational example for the diffusion equation with boundary control and point observation

一个新的ν-度量计算示例用于带有边界控制和点观测的扩散方程

Amol Sasane

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出了一种新的ν-度量计算方法,用于解决具有参数不确定性的扩散方程控制问题,通过边界控制和点观测实现鲁棒性分析。

Comments 11 pages, 2 figues

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11345 2025-12-15 cs.LG cs.RO 78%

Symmetry-Aware Steering of Equivariant Diffusion Policies: Benefits and Limits

具有对称性的扩散策略引导:益处与局限

Minwoo Park, Junwoo Chang, Jongeun Choi, Roberto Horowitz

机构 * Yonsei University(延世大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出了一种具有对称性的扩散策略引导框架,通过利用对称性提升样本效率和策略性能,同时揭示了严格等变的实践边界。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08929 2025-12-15 math.AP 78%

On a cross-diffusion hybrid model: Cancer Invasion Tissue with Normal Cell Involved

关于交叉扩散混合模型:涉及正常细胞的肿瘤侵袭组织

Guanjun Pan, Hong-Ming Yin

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出了一种涉及正常细胞的肿瘤侵袭模型,通过交叉扩散混合系统微分方程研究肿瘤侵袭的数学适定性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02757 2025-12-15 eess.SP 78%

Channel Knowledge Map Construction via Physics-Inspired Diffusion Model Without Prior Observations

通过物理启发的扩散模型构建通道知识图谱而不依赖先验观测

Yunzhe Zhu, Xuewen Liao, Zhenzhen Gao, Linzhou Zeng, Yong Zeng

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出了一种无需先验观测的物理启发扩散模型,用于构建高精度的6G无线系统通道知识图谱,通过整合物理约束提升生成结果的准确性和物理一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20445 2025-12-15 cs.LG physics.plasm-ph 78%

Diffusion for Fusion: Designing Stellarators with Generative AI

扩散用于融合:利用生成式AI设计星形装置

Misha Padidar, Teresa Huang, Andrew Giuliani, Marina Spivak

机构 * Center for Computational Mathematics(计算数学中心) Flatiron Institute(Flatiron研究所)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出利用生成式AI快速生成具有特定特性的星形装置设计,通过条件扩散模型在QUASR数据库上训练,生成准对称星形装置并展示其性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18078 2025-12-15 cs.SD cs.AI 78%

Diffusion-based Surrogate Model for Time-varying Underwater Acoustic Channels

基于扩散模型的时变水下声学信道代理模型

Kexin Li, Mandar Chitre

机构 * ARL, Tropical Marine Science Institute, National University of Singapore(ARL、热带海洋科学研究所、新加坡国立大学) Department of Electrical and Computer Engineering, National University of Singapore(电气与计算机工程系、新加坡国立大学)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出StableUASim,一种基于扩散模型的预训练代理模型,用于高效建模水下声学信道,实现快速适应和高精度仿真。

Comments Updated references with DOIs

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14040 2025-12-15 cs.LG cs.AI cs.RO 78%

WARPD: World model Assisted Reactive Policy Diffusion

WARPD:世界模型辅助的反应策略扩散

Shashank Hegde, Satyajeet Das, Gautam Salhotra, Gaurav S. Sukhatme

机构 * University of Southern California(南加州大学) Google(谷歌) Intrinsic LLC(Intrinsic 公司)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 WARPD通过直接生成闭环策略,提高了机器人任务中长动作时间跨度和鲁棒性的性能,同时显著降低了推理成本。

Comments Outstanding Paper Award at the Embodied World Models for Decision Making Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01220 2025-12-15 cs.CV cs.AI 74%

Tera-MIND: Tera-scale mouse brain simulation via spatial mRNA-guided diffusion

Tera-MIND:通过空间mRNA引导扩散实现千兆级小鼠大脑模拟

Jiqing Wu, Ingrid Berg, Yawei Li, Ender Konukoglu, Viktor H. Koelzer

机构 * Department of Biomedical Engineering, University of Basel, Switzerland(巴塞尔大学生物医学工程系) Department of Pathology and Molecular Pathology, University Hospital, University of Zurich, Switzerland(苏黎世大学病理学与分子病理学系,苏黎世大学医院) Computer Vision Lab, ETH Zurich, Switzerland(苏黎世联邦理工学院计算机视觉实验室) Integrated System Laboratory, ETH Zurich, Switzerland(苏黎世联邦理工学院集成系统实验室) Institute of Medical Genetics and Pathology, University Hospital Basel, Switzerland(巴塞尔大学医学遗传学与病理学研究所,巴塞尔大学医院)

专题命中 扩散模型 :diffusion(title);分类 cs.CV

AI总结 Tera-MIND通过空间mRNA引导扩散模型,实现千兆级小鼠大脑的三维模拟,用于研究大脑分子相互作用和生物医学应用

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11111 2025-12-15 math.NA cs.NA 71%

Analysis of a Discontinuous Galerkin Method for Diffusion Problems on Intersecting Domains

对交叠域上扩散问题的不连续伽辽金方法分析

Miroslav Kuchta, Rami Masri, Beatrice Riviere

专题命中 扩散模型 :diffusion(title)

AI总结 本文分析了在交叠域上求解扩散问题的不连续伽辽金方法,证明了稳定性与收敛性,并通过数值实验验证了理论结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18617 2025-12-15 nlin.CD cond-mat.stat-mech math-ph math.MP 71%

Diffusion as a Signature of Chaos

扩散作为混沌的特征

Nachiket Karve, Nathan Rose, David Campbell

专题命中 扩散模型 :diffusion(title)

AI总结 通过引入可观测漂移,研究揭示了经典混沌与抗性变形敏感性的等价性,并通过数值实验验证了这一结论。

Comments 17 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11225 2025-12-15 cs.CV cs.AI cs.LG 70%

VFMF: World Modeling by Forecasting Vision Foundation Model Features

VFMF:通过预测视觉基础模型特征进行世界建模

Gabrijel Boduljak, Yushi Lan, Christian Rupprecht, Andrea Vedaldi

专题命中 扩散模型 :image generation(abstract);diffusion(abstract);分类 cs.CV

AI总结 本文提出了一种基于VFM特征的生成预测器,通过自回归流匹配在潜在空间中实现更准确的预测,提升世界建模的效率和效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02069 2025-12-15 cs.GR cs.CV 62%

Spec-Gloss Surfels and Normal-Diffuse Priors for Relightable Glossy Objects

基于Spec-Gloss Surfels和正常-漫反射先验的可重照明光泽物体

Georgios Kouros, Minye Wu, Tinne Tuytelaars

机构 * Department of Electrical Engineering (ESAT), KU Leuven, Belgium(电气工程系(ESAT),比利时鲁文大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV、cs.GR

AI总结 本文提出了一种基于Spec-Gloss Surfels和正常-漫反射先验的可重照明框架,通过整合微面BRDF与镜面-光泽参数化,提升光泽物体的重建和重照明质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11792 2025-12-15 cs.CV 57%

Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation

从跟踪中获取结构:从自回归视频跟踪模型中蒸馏保持结构的运动用于视频生成

Yang Fei, George Stoica, Jingyuan Liu, Qifeng Chen, Ranjay Krishna, Xiaojuan Wang, Benlin Liu

机构 * HKUST(香港科技大学) University of Washington(华盛顿大学) Georgia Tech(佐治亚理工学院) Adobe(Adobe公司)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 通过蒸馏自回归视频跟踪模型中的结构保持运动先验,SAM2VideoX在视频生成任务中实现了更高的保真度和一致性。

Comments Project Website: https://sam2videox.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11719 2025-12-15 cs.CV 57%

Referring Change Detection in Remote Sensing Imagery

遥感图像中的指称变化检测

Yilmaz Korkmaz, Jay N. Paranjape, Celso M. de Melo, Vishal M. Patel

机构 * Johns Hopkins University(约翰霍普金斯大学) DEVCOM U.S. Army Research Laboratory(美国陆军研究实验室)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出指称变化检测方法,通过自然语言提示实现遥感图像中特定类别的变化检测,并引入两阶段框架提升数据生成效率。

Comments 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏