arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-12-15 至 2025-12-15 共收录 61 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 47 篇

2512.11284 2025-12-15 cs.CV 57%

RcAE: Recursive Reconstruction Framework for Unsupervised Industrial Anomaly Detection

RcAE:递归重建框架用于无监督工业异常检测

Rongcheng Wu, Hao Zhu, Shiying Zhang, Mingzhe Wang, Zhidong Li, Hui Li, Jianlong Zhou, Jiangtao Cui, Fang Chen, Pingyang Sun, Qiyu Liao, Ye Lin

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 RcAE通过递归重建框架有效检测工业异常,结合CRD模块和DPN网络,在参数更少的情况下实现与扩散模型相当的性能。

Comments 19 pages, 7 figures, to be published in AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11253 2025-12-15 cs.CV 57%

PersonaLive! Expressive Portrait Image Animation for Live Streaming

PersonaLive! 用于直播的 expressive 人物图像动画

Zhiyuan Li, Chi-Man Pun, Chen Fang, Jue Wang, Xiaodong Cun

机构 * University of Macau(澳门大学) GVC Lab, Great Bay University(大湾大学GVC实验室)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 PersonaLive通过多阶段训练方法实现了实时面部动画生成,提升了效率和稳定性,达到7-22倍的速度提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11199 2025-12-15 cs.CV cs.LG 57%

CADKnitter: Compositional CAD Generation from Text and Geometry Guidance

CADKnitter:从文本和几何引导生成组合CAD模型

Tri Le, Khang Nguyen, Baoru Huang, Tung D. Ta, Anh Nguyen

机构 * FPT Software AI Center(FPT软件AI中心) Department of Robotics, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(机器人系,Mohamed bin Zayed人工智能大学) Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) Department of Creative Informatics, The University of Tokyo(创意信息学系,东京大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 CADKnitter通过几何引导扩散策略生成组合CAD模型,结合文本和几何约束,提升多部件装配设计效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16979 2025-12-15 cs.CV cs.AI 57%

The Finer the Better: Towards Granular-aware Open-set Domain Generalization

更细的更好:面向粒度感知的开集域泛化

Yunyun Wang, Zheng Duan, Xinyue Liao, Ke-Jia Chen, Songcan Chen

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 SeeCLIP通过细粒度语义增强解决开集域泛化中的结构风险与开放空间风险困境,提升模型对难区分未知类别的判别能力。

Comments 9 pages,3 figures,aaai2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05046 2025-12-15 cs.CV 57%

FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing

FlowDirector: 无需训练的流引导文本到视频编辑

Guangzhao Li, Yanming Yang, Chenxi Song, Chi Zhang

机构 * AGI Lab, Westlake University(西湖大学AGI实验室) Central South University(中南大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 FlowDirector通过直接在数据空间中建模编辑过程,无需训练和倒置步骤,实现了更精确的文本到视频编辑。

Comments Project Page is https://flowdirector-edit.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08518 2025-12-15 cs.CV 57%

Visual-Friendly Concept Protection via Selective Adversarial Perturbations

通过选择性对抗扰动实现视觉友好的概念保护

Xiaoyue Mi, Fan Tang, You Wu, Juan Cao, Peng Li, Yang Liu

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出VCPro框架,通过低感知性对抗扰动实现图像关键概念的保护,平衡扰动可见性与保护效果。

Comments AAAI AISI 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01004 2025-12-15 cs.CV cs.AI 57%

MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing

MoCA-Video:面向一致视频编辑的运动感知概念对齐

Tong Zhang, Juan C Leon Alcazar, Victor Escorcia, Bernard Ghanem

机构 * Department of Computer Science(计算机科学系) King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学科学与技术学院)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 MoCA-Video通过在冻结的视频扩散模型潜在空间中进行语义混合,实现无需训练的高质量视频编辑,通过动量修正和伽马残差模块提升时间稳定性和语义一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11525 2025-12-15 cs.LG cs.AI 50%

NeuralOGCM: Differentiable Ocean Modeling with Learnable Physics

NeuralOGCM:可微海洋建模与可学习物理

Hao Wu, Yuan Gao, Fan Xu, Fan Zhang, Guangliang Liu, Yuxuan Liang, Xiaomeng Huang

机构 * Tsinghua University, Beijing, China(清华大学) University of Science and Technology of China, Hefei, China(中国科学技术大学) The Chinese University of Hong Kong, Hong Kong, China(香港中文大学) Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China(香港科学与技术大学(广州))

专题命中 扩散模型 :diffusion(abstract)

AI总结 NeuralOGCM通过融合可微编程与深度学习,实现高效且物理合理的海洋建模,提升科学计算的精度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11489 2025-12-15 math.AP 50%

Effective transmission through an interface with evolving microstructure

通过具有演化的微观结构界面的有效传输

Lucas M. Fix, Gianna Götzmann, Malte A. Peter, Jan-F. Pietschmann

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文研究了具有演化微观结构界面的非线性反应-扩散-对流方程组的有效模型,通过双重收敛和展开技术推导出在 $\varepsilon \to 0$ 极限下的有效模型。

Comments 44 pages, 2 figures, comments welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11370 2025-12-15 math.AP 50%

A variational approach to nonlocal heat equations

变分方法用于非局部热方程

Edoardo Mainini

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文提出了一种变分方法用于非局部热方程,通过加权积分方法解决非局部线性扩散模型中的强迫项问题,并提供了解的选择原理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11244 2025-12-15 eess.SY cs.SY q-bio.BM 50%

Model Reduction of Multicellular Communication Systems via Singular Perturbation: Sender Receiver Systems

通过奇异性扰动方法进行多细胞通信系统模型简化:发件人接收系统

Taishi Kotsuka, Enoch Yeung

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文通过奇异性扰动方法简化多细胞通信系统模型,利用闭合形式通信矩阵实现大规模细胞群体的高效模拟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05486 2025-12-15 math.NA cs.NA 50%

A New Class of General Linear Method with Inherent Quadratic Stability for Solving Stiff Differential Systems

一种具有内在二次稳定性的新一类通用线性方法用于求解刚性微分系统

Sakshi Gautam, Ram K. Pandey

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文提出了一种具有内在二次稳定性的新通用线性方法,用于求解刚性微分系统,通过阶条件和稳定性约束构建,并在多个实际问题中验证其有效性。

Comments 19 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23181 2025-12-15 cond-mat.mtrl-sci cs.LG 50%

Introducing physics-informed generative models for targeting structural novelty in the exploration of chemical space

引入物理指导的生成模型以针对化学空间中的结构新颖性

Andrij Vasylenko, Federico Ottomano, Christopher M. Collins, Rahul Savani, Matthew S. Dyer, Matthew J. Rosseinsky

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文提出物理指导的生成模型,用于在化学空间中寻找结构新颖的材料,通过结合局部环境多样性和紧凑性来平衡物理合理性和新颖性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08187 2025-12-15 cond-mat.soft cond-mat.stat-mech 50%

Stability of discrete-symmetry flocks: sandwich state, traveling domains and motility-induced pinning

离散对称性群的稳定性:夹层状态、行进域和运动诱导钉定

Swarnajit Chatterjee, Mintu Karmakar, Matthieu Mangeat, Heiko Rieger, Raja Paul

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文研究了离散活性物质中群的稳定性,发现小滴状物可诱导夹层状态和相反转,揭示了运动诱导钉定现象及多态共存机制。

Comments 19 pages, 19 figures

Journal ref Phys. Rev. E 112, 064115 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00731 2025-12-15 cs.LG cs.AI 50%

iPINNER: An Iterative Physics-Informed Neural Network with Ensemble Kalman Filter

iPINNER: 一种基于集合卡尔曼滤波的迭代物理信息神经网络

Binghang Lu, Changhong Mou, Guang Lin

机构 * School of Electrical and Computer Engineering(电气与计算机工程学院) Purdue University(普渡大学) Department of Mathematics(数学系) Department of Mathematics and School of Mechanical Engineering(数学系和机械工程学院)

专题命中 扩散模型 :diffusion(abstract)

AI总结 iPINNER通过集合卡尔曼滤波和NSGA-III优化,提升PINNs在正向和逆问题中处理噪声数据和缺失物理的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14384 2025-12-15 cond-mat.stat-mech cond-mat.soft cond-mat.str-el 50%

Universal scaling in one-dimensional non-reciprocal matter

一维非 reciprocal 物质中的普遍标度

Shuoguang Liu, Peter B. Littlewood, Ryo Hanai

专题命中 扩散模型 :diffusion(abstract)

AI总结 研究一维非 reciprocal 物质中的普遍标度行为,揭示非平衡临界现象及异常粗糙指数。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07741 2025-12-15 cond-mat.stat-mech math-ph math.MP nlin.CG 50%

Critical behavior of a programmable time-crystal lattice gas

可编程时间晶体晶格气体的临界行为

R. Hurtado-Gutiérrez, C. Pérez-Espigares, P. I. Hurtado

专题命中 扩散模型 :diffusion(abstract)

AI总结 研究通过时间晶体晶格气体模型揭示了可编程时间晶体的临界行为,利用蒙特卡洛模拟分析了不同阶数下的非平衡相变特性及临界指数。

Comments 14 pages, 8 figures. arXiv admin note: text overlap with arXiv:2404.04135, arXiv:2406.08581

Journal ref Phys. Rev. E 112, 044135 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14973 2025-12-15 cond-mat.soft cond-mat.mtrl-sci physics.optics 50%

Expanding the reach of diffusing wave spectroscopy and tracer bead microrheology

扩展扩散波光谱和示踪珠微流变学的应用范围

Manuel Helfer, Chi Zhang, Frank Scheffold

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文提出改进的双细胞回声DWS方案,通过无需校准的数据融合和指数基底拟合,提升微流变学数据质量,解决短相关时间准确性问题。

Comments 11 pages, 5 figures

Journal ref Phys. Rev. Research 7, 043274 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01933 2025-12-15 cs.LG 50%

TAEGAN: Generating Synthetic Tabular Data For Data Augmentation

TAEGAN:生成合成表格数据用于数据增强

Jiayu Li, Zilong Zhao, Kevin Yee, Uzair Javaid, Biplab Sikdar

机构 * National University of Singapore(国立新加坡大学) Betterdata AI Singapore(Betterdata AI新加坡)

专题命中 扩散模型 :diffusion(abstract)

AI总结 TAEGAN通过引入自监督预热训练和改进的损失函数,提升表格数据生成的稳定性与效率,优于现有方法。

Comments This paper is accepted at ACML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15361 2025-12-15 math.AP 50%

A vector-host epidemic model with spatial structure and seasonality

具有空间结构和季节性的向量-宿主流行病模型

Mingxin Wang, Qianying Zhang

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文研究了具有空间结构和季节性的向量-宿主流行病模型,采用上界和下界解方法分析周期反应-扩散模型的阈值参数R0。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11180 2025-12-15 q-bio.PE 50%

Optimal Protected Area Design for Allee Effect Mitigation in Spatial Predator-Prey Systems

在空间捕食者-猎物系统中通过阿莱效应缓解设计最优保护区域

Junhui Hu

专题命中 扩散模型 :diffusion(abstract)

AI总结 本研究通过设计最优保护区域,缓解空间捕食者-猎物系统中阿莱效应导致的低密度种群灭绝问题,采用双目标优化方法分析保护区域配置对种群恢复的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10991 2025-12-15 cs.LG cs.AI physics.chem-ph q-bio.QM 50%

MolSculpt: Sculpting 3D Molecular Geometries from Chemical Syntax

MolSculpt:从化学语法雕刻3D分子几何

Zhanpeng Chen, Weihao Gao, Shunyu Wang, Yanan Zhu, Hong Meng, Yuexian Zou

机构 * AI for Science (AI4S)-Preferred Program, Peking University Shenzhen Graduate School, China(人工智能科学(AI4S)优选计划,北京大学深圳研究生院,中国) Faculty of Materials Science, Shenzhen MSU-BIT University, Shenzhen, China(材料科学学院,深圳MSU-BIT大学,中国)

专题命中 扩散模型 :diffusion(abstract)

AI总结 MolSculpt通过整合1D化学知识与3D生成过程,实现了从化学语法到精确3D分子几何的高效生成。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 可控生成 6 篇

2512.11720 2025-12-15 cs.CV 83%

Reframing Music-Driven 2D Dance Pose Generation as Multi-Channel Image Generation

将音乐驱动的2D舞蹈姿态生成重新框架化为多通道图像生成

Yan Zhang, Han Zou, Lincong Feng, Cong Xie, Ruiqi Yu, Zhenpeng Zhan

机构 * Global Business Unit, Baidu Inc(百度公司全球业务部)

专题命中 可控生成 :image generation(title);text-to-image(abstract);image synthesis(abstract);分类 cs.CV

AI总结 将音乐驱动的2D舞蹈姿态生成重新框架化为多通道图像生成,通过时间共享索引和参考姿态条件策略提升生成质量与一致性

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00996 2025-12-15 cs.CV 79%

Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models

具有时间推理的上下文微调用于视频扩散模型的多功能控制

Kinam Kim, Junha Hyung, Jaegul Choo

机构 * KAIST AI(韩国科学技术院人工智能研究中心)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 TIC-FT通过时间推理实现视频扩散模型的多功能控制,无需架构修改,仅需少量样本即可实现高质量视频生成。

Comments project page: https://kinam0252.github.io/TIC-FT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06215 2025-12-15 cs.CV 70%

Fine-grained Defocus Blur Control for Generative Image Models

细粒度模糊模糊控制用于生成图像模型

Ayush Shrivastava, Connelly Barnes, Xuaner Zhang, Lingzhi Zhang, Andrew Owens, Sohrab Amirghodsi, Eli Shechtman

机构 * University of Michigan(密歇根大学) Adobe Research(Adobe研究)

专题命中 可控生成 :text-to-image(abstract);diffusion(abstract);分类 cs.CV

AI总结 本文提出了一种基于EXIF数据的文本到图像扩散模型,通过生成可控的镜头模糊效果,实现细粒度的模糊控制,提升图像生成的可控性与真实性。

Comments Project link: https://www.ayshrv.com/defocus-blur-gen

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11645 2025-12-15 cs.CV 57%

FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint

FactorPortrait: 通过解耦的表情、姿态和视角实现可控的肖像动画

Jiapeng Tang, Kai Li, Chengxiang Yin, Liuhao Ge, Fei Jiang, Jiu Xu, Matthias Nießner, Christian Häne, Timur Bagautdinov, Egor Zakharov, Peihong Guo

机构 * Meta Reality Labs(Meta现实实验室) Technical University of Munich(慕尼黑技术大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

AI总结 FactorPortrait通过解耦表情、姿态和视角实现可控肖像动画,利用预训练编码器和3D网格跟踪生成逼真动态效果。

Comments Project page: https://tangjiapeng.github.io/FactorPortrait/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09585 2025-12-15 cs.SD cs.MM 57%

Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation

视频中的音乐:语义、时间与节奏对齐用于视频到音乐生成

Xinyi Tong, Yiran Zhu, Jishang Chen, Chunru Zhan, Tianle Wang, Sirui Zhang, Nian Liu, Tiezheng Ge, Duo Xu, Xin Jin, Feng Yu, Song-Chun Zhu

专题命中 可控生成 :diffusion(abstract);分类 cs.MM

AI总结 VeM通过语义、时间和节奏对齐,生成高质量视频到音乐,提升视听沉浸感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05029 2025-12-15 cs.CV 57%

Estimating Object Physical Properties from RGB-D Vision and Depth Robot Sensors Using Deep Learning

基于RGB-D视觉和深度机器人传感器利用深度学习估计物体物理特性

Ricardo Cardoso, Plinio Moreno

机构 * Ricardo Pedreiras Cardoso Plinio Moreno

专题命中 可控生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出结合深度图像和RGB图像的数据,利用深度学习方法估计物体质量,通过合成数据集提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 图像修复 3 篇

2506.14541 2025-12-15 cs.CV 83%

Exploring Diffusion with Test-Time Training on Efficient Image Restoration

探索通过测试时间训练的高效图像恢复

Rongchang Lu, Tianduo Luo, Yunzhi Jiang, Conghan Yue, Pei Yang, Guibao Liu, Changyang Gu

专题命中 图像修复 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 DiffRWKVIR通过结合测试时间训练与高效扩散,实现了高效的图像恢复,在多个基准测试中表现出色。

Comments We withdraw this paper due to erroneous experiment data in the ablation study, which was inadvertently copied from our preprint "Ultra-Lightweight Semantic-Injected Imagery Super-Resolution for Real-Time UAV Remote Sensing" This nearly constituted academic misconduct. We sincerely apologize and thank those who alerted us

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11640 2025-12-15 cs.CV 57%

COSMO-INR: Complex Sinusoidal Modulation for Implicit Neural Representations

COSMO-INR:复数正弦调制用于隐式神经表示

Pandula Thennakoon, Avishka Ranasinghe, Mario De Silva, Buwaneka Epakanda, Roshan Godaliyadda, Parakrama Ekanayake, Vijitha Herath

机构 * Department of Electrical and Electronic Engineering(电气与电子工程系)

专题命中 图像修复 :inpainting(abstract);分类 cs.CV

AI总结 COSMO-INR通过复数正弦调制激活函数提升隐式神经表示的性能,实现图像重建、去噪、超分辨率等任务的显著改进。

Comments Submitted as a conference paper to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏