arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-12-16 至 2025-12-16 共收录 116 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 86 篇

2403.14930 2025-12-16 cond-mat.soft physics.comp-ph 50%

A numerical framework for phoretic particles

用于非牛顿粒子的动力学的数值框架

Zhe Gou, Alexander Farutin, Chaouqi Misbah

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文提出了一种用于研究非牛顿粒子动力学的数值框架,通过边界积分方法和重叠网格方法实现了高精度的流体-结构相互作用模拟。

Comments 26 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12129 2025-12-16 cs.SD 50%

A comparative study of generative models for child voice conversion

生成模型在儿童语音转换中的比较研究

Protima Nomo Sudro, Anton Ragni, Thomas Hain

机构 * Department of Computer Science, The University of Sheffield(谢菲尔德大学计算机科学系)

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文比较了四种生成模型在成人到儿童语音转换中的性能,提出了一种高效频率扭曲技术以提升转换质量。

Comments 6 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11912 2025-12-16 cs.AI 50%

Robustness of Probabilistic Models to Low-Quality Data: A Multi-Perspective Analysis

概率模型对低质量数据的鲁棒性:多视角分析

Liu Peng, Yaochu Jin

机构 * Trustworthy and General AI Lab(可信与通用人工智能实验室) Department of Artificial Intelligence(人工智能系)

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文通过多视角分析,揭示了概率模型在低质量数据下的鲁棒性差异,发现自回归模型更具韧性,而扩散模型表现较差,主要受条件信息丰富性和训练数据信息含量影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11828 2025-12-16 physics.ao-ph physics.comp-ph physics.flu-dyn 50%

Numerical Assessment of Advective and Diffusive Dynamics of Interacting and Isolated Prototypical Convectively Initiated Circulations

对相互作用和孤立典型对流性循环的辐合与扩散动力学的数值评估

Matthew R. Igel, Joseph A. Biello, Adele L. Igel

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文通过KRoNUT模型评估了相互作用和孤立对流环流的辐合与扩散动力学,揭示了辐合和扩散在环流演变中的作用及相互作用环流的动力学多样性。

Comments Companion article to "Idealized Cumulus Cloud-Scale Motions and the Dynamics of Isolated and Coupled Flows"

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10516 2025-12-16 astro-ph.GA astro-ph.SR 50%

Isotopomer-Specific Carbon Isotope Ratio of Complex Organic Molecules in Star-Forming Cores

恒星形成核心中复杂有机分子的同位素omer特异性碳同位素比值

Ryota Ichimura, Hideko Nomura, Kenji Furuya, Tetsuya Hama, T. J. Millar

专题命中 扩散模型 :diffusion(abstract)

AI总结 本研究通过构建新的天体化学反应网络,解析恒星形成核心中复杂有机分子的同位素omer特异性碳同位素比值,揭示其形成路径与环境。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21339 2025-12-16 cond-mat.mtrl-sci 50%

Quantum Monte Carlo Benchmarking of Molecular Adsorption on Graphene-Supported Single Pt Atom

石墨烯支撑单个铂原子上分子吸附的量子蒙特卡洛基准测试

Jeonghwan Ahn, Iuegyun Hong, Gwangyoung Lee, Hyeondeok Shin, Anouar Benali, Yongkyung Kwon

专题命中 扩散模型 :diffusion(abstract)

AI总结 本研究通过量子蒙特卡洛方法评估了石墨烯支撑单个铂原子上分子吸附的计算性能,揭示了DFT与DMC在吸附能预测上的差异及CO中毒问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01530 2025-12-16 cs.SD eess.AS stat.AP 50%

Generative AI-based data augmentation for improved bioacoustic classification in noisy environments

基于生成AI的数据增强以改进噪声环境中的生物声学分类

Anthony Gibbons, Emma King, Ian Donohue, Andrew Parnell

专题命中 扩散模型 :diffusion(abstract)

AI总结 本研究利用生成AI模型合成音频频谱图,通过数据增强提升噪声环境下生物声学分类的准确性,展示了在稀有物种检测中的应用潜力。

Comments 25 pages, 4 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 可控生成 7 篇

2510.09561 2025-12-16 cs.CV 79%

TC-LoRA: Temporally Modulated Conditional LoRA for Adaptive Diffusion Control

TC-LoRA: 基于时间调节的条件LoRA用于自适应扩散控制

Minkyoung Cho, Ruben Ohana, Christian Jacobsen, Adityan Jothi, Min-Hung Chen, Z. Morley Mao, Ethem Can

机构 * University of Michigan(密歇根大学) NVIDIA(英伟达)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 TC-LoRA通过动态调整模型权重实现自适应扩散控制,提升生成保真度和空间条件符合性。

Comments Project Page: https://minkyoungcho.github.io/tc-lora/; NeurIPS 2025 Workshop on SPACE in Vision, Language, and Embodied AI (SpaVLE); 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17364 2025-12-16 cs.CV cs.AI 79%

Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation

条件编织与专家调节:迈向通用且可控的图像生成

Guoqing Zhang, Xingtong Ge, Lu Shi, Xin Zhang, Muqing Xue, Wanru Xu, Yigang Cen, Yidong Li

机构 * State Key Laboratory of Advanced Rail Autonomous Operation(先进轨道交通自主运行国家重点实验室) School of Computer Science and Technology(计算机科学与技术学院) Visual Intellgence +X International Cooperation Joint Laboratory of MOE(教育部视觉智能+X国际合作联合实验室) Hong Kong University of Science and Technology(香港科技大学) SenseTime Research(商汤科技研究院) Beijing Jiaotong University(北京交通大学) SenseTime Research Institute(时光机器研究院)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

AI总结 提出UniGen框架,通过CoMoE模块和WeaveNet机制实现通用且可控的图像生成,提升效率和表达性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00227 2025-12-16 cs.CV cs.AI cs.RO 79%

Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes

Ctrl-Crash: 可控扩散用于逼真汽车碰撞

Anthony Gosselin, Ge Ya Luo, Luis Lara, Florian Golemo, Derek Nowrouzezahrai, Liam Paull, Alexia Jolicoeur-Martineau, Christopher Pal

机构 * McGill University(麦吉尔大学) CIFAR AI Chair(CIFAR人工智能 chair)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 Ctrl-Crash通过可控扩散生成逼真汽车碰撞视频,提升交通安全模拟的可控性和真实性。

Comments Under review at Pattern Recognition Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13014 2025-12-16 cs.CV 70%

JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation Promotion

JoDiffusion: 通过像素级标注联合扩散图像以促进语义分割

Haoyu Wang, Lei Zhang, Wenrui Liu, Dengyang Jiang, Wei Wei, Chen Ding

专题命中 可控生成 :image generation(abstract);diffusion(abstract);分类 cs.CV

AI总结 JoDiffusion通过联合扩散图像与像素级标注,提升语义分割的性能和可扩展性。

Comments Accepted at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13008 2025-12-16 cs.CV 57%

TWLR: Text-Guided Weakly-Supervised Lesion Localization and Severity Regression for Explainable Diabetic Retinopathy Grading

TWLR: 基于文本引导的弱监督病变定位与严重程度回归用于可解释性糖尿病视网膜病变分级

Xi Luo, Shixin Xu, Ying Xie, JianZhong Hu, Yuwei He, Yuhui Deng, Huaxiong Huang

机构 * Guangdong Provincial Key Laboratory of Interdisciplinary Research and Application for Data Science(广东省级交叉学科研究与数据科学应用重点实验室) Department of Statistics and Data Science, Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist大学统计与数据科学系) Faculty of Science, Hong Kong Baptist University(香港 Baptist大学科学学院) Data Science Research Center, Duke Kunshan University(杜克-昆山大学数据科学研究中心) Shanxi Provincial People’s Hospital(山西人民医院) The Fifth Clinical Medical school of Shanxi Medical University(山西医科大学第五临床医学院) Research Center for Mathematics, Beijing Normal University(北京师范大学数学研究中心) Department of Mathematics and Statistics, York University(约克大学数学与统计学系)

专题命中 可控生成 :inpainting(abstract);分类 cs.CV

AI总结 TWLR通过双阶段框架实现糖尿病视网膜病变的可解释性评估,结合视觉语言模型和弱监督分割,实现病变定位与严重程度回归。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12751 2025-12-16 cs.CV 57%

GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation

GenieDrive: 向具有物理意识的驾驶世界模型迈进:基于4D占用的视频生成

Zhenya Yang, Zhe Liu, Yuxiang Lu, Liping Hou, Chenxuan Miao, Siyi Peng, Bailan Feng, Xiang Bai, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Huazhong University of Science and Technology(华中科技大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

AI总结 GenieDrive通过4D占用引导的视频生成,实现物理意识的驾驶视频生成,提升预测精度和视频质量。

Comments The project page is available at https://huster-yzy.github.io/geniedrive_project_page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12664 2025-12-16 cs.CV 57%

InteracTalker: Prompt-Based Human-Object Interaction with Co-Speech Gesture Generation

InteracTalker: 基于提示的人-物体交互与同步手势生成

Sreehari Rajan, Kunal Bhosikar, Charu Sharma

机构 * Machine Learning Lab, IIIT Hyderabad, India(IIIT Hyderabad 机器学习实验室)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

AI总结 InteracTalker通过整合基于提示的物体感知交互与同步手势生成,实现了人-物体交互的统一框架,提升了动作的真实性和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 图像修复 2 篇

2512.12060 2025-12-16 cs.CV cs.LG cs.MM 81%

CreativeVR: Diffusion-Prior-Guided Approach for Structure and Motion Restoration in Generative and Real Videos

CreativeVR: 一种基于扩散先验的结构与运动修复方法用于生成视频和真实视频

Tejas Panambur, Ishan Rajendrakumar Dave, Chongjian Ge, Ersin Yumer, Xue Bai

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Adobe(Adobe公司)

专题命中 图像修复 :diffusion(title,abstract);分类 cs.CV、cs.MM

AI总结 CreativeVR通过扩散先验引导方法,有效修复AI生成和真实视频中的结构与运动伪影,实现高质量视频修复。

Comments The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04130 2025-12-16 cs.CV eess.IV physics.optics 57%

Deep priors for satellite image restoration with accurate uncertainties

具有准确不确定性的卫星图像修复中的深度先验

Biquard Maud, Marie Chabert, Florence Genin, Christophe Latry, Thomas Oberlin

机构 * ISAE-Supaero / CNES(ISAE-苏帕欧 / CNES) IRIT/INP-ENSEEIHT CNES ISAE-Supaero(ISAE-苏帕欧)

专题命中 图像修复 :diffusion(abstract);分类 cs.CV

AI总结 本文提出VBLE-xz方法,通过变分压缩自动编码器的潜在空间解决卫星图像修复的反问题,并联合估计潜在空间和图像空间的不确定性,为需要不确定性量化的场景提供了一种高效替代方案。

Journal ref IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1-16, 2025, Art no. 5652916

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 个性化与一致性 3 篇

2512.12963 2025-12-16 cs.CV 79%

SCAdapter: Content-Style Disentanglement for Diffusion Style Transfer

SCAdapter: 内容-风格解耦用于扩散风格迁移

Luan Thanh Trinh, Kenji Doi, Atsuki Osanai

机构 * LY Corporation(LY公司)

专题命中 个性化与一致性 :diffusion(title,abstract);分类 cs.CV

AI总结 SCAdapter通过CLIP图像空间实现内容与风格的解耦,提升扩散模型在逼真图像迁移中的效果和效率。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13665 2025-12-16 cs.CV 57%

Grab-3D: Detecting AI-Generated Videos from 3D Geometric Temporal Consistency

Grab-3D: 通过3D几何时间一致性检测AI生成视频

Wenhan Chen, Sezer Karaoglu, Theo Gevers

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 个性化与一致性 :diffusion(abstract);分类 cs.CV

AI总结 Grab-3D通过3D几何时间一致性检测AI生成视频,利用几何感知Transformer框架提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03738 2025-12-16 cs.CV 57%

FACM: Flow-Anchored Consistency Models

FACM:基于流的一致性模型

Yansong Peng, Kai Zhu, Yu Liu, Pingyu Wu, Hebei Li, Xiaoyan Sun, Feng Wu

机构 * University of Science and Technology of China(中国科学技术大学) Tongyi Lab(通义实验室)

专题命中 个性化与一致性 :text-to-image(abstract);分类 cs.CV

AI总结 FACM通过流锚定方法解决连续时间一致性模型的训练不稳定性问题,实现高效生成和稳定训练。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 图像生成评测 1 篇

2512.12287 2025-12-16 cs.CV 57%

RealDrag: The First Dragging Benchmark with Real Target Image

RealDrag: 首个包含真实目标图像的拖动基准

Ahmad Zafarani, Zahra Dehghanian, Mohammadreza Davoodi, Mohsen Shadroo, MohammadAmin Fazli, Hamid R. Rabiee

机构 * Sharif University of Technology(谢里夫技术大学) Iran University of Science and Technology(伊朗科学技术大学) Comprehensive University of Islamic Revolution(伊斯兰革命综合大学)

专题命中 图像生成评测 :image editing(abstract);分类 cs.CV

AI总结 RealDrag是首个包含真实目标图像的拖动基准,通过提出四个任务特定指标,系统评估了17种最先进模型,揭示了当前方法间的权衡并建立了可重复的基线。

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 效率与蒸馏 6 篇

2512.13006 2025-12-16 cs.CV 89%

Few-Step Distillation for Text-to-Image Generation: A Practical Guide

文本到图像生成的少步蒸馏:一份实用指南

Yifan Pu, Yizeng Han, Zhiwei Tang, Jiasheng Tang, Fan Wang, Bohan Zhuang, Gao Huang

机构 * Tsinghua University(清华大学) DAMO Academy, Alibaba Group(阿里达摩院) Hupan Lab(华研实验室) Zhejiang University(浙江大学)

专题命中 效率与蒸馏 :text-to-image(title,abstract);image generation(title);diffusion(abstract);image synthesis(abstract)

AI总结 本文提出了一种少步蒸馏方法,用于文本到图像生成,通过比较先进技术并提供实用指南,以提高生成质量和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12889 2025-12-16 cs.LG 78%

Distillation of Discrete Diffusion by Exact Conditional Distribution Matching

通过精确条件分布匹配实现离散扩散模型的蒸馏

Yansong Gao, Yu Sun

机构 * Google(谷歌)

专题命中 效率与蒸馏 :diffusion(title,abstract)

AI总结 本文提出了一种基于条件分布匹配的蒸馏方法,用于加速离散扩散模型的推断过程,通过匹配教师和学生模型的条件分布来提高效率。

Comments [work in progress]

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18461 2025-12-16 cs.CV 57%

An Efficient and Harmonized Framework for Balanced Cross-Domain Feature Integration

一种高效且协调的平衡跨域特征整合框架

Shaoxu Li, Ye Pan

机构 * Shaoxu Li, Ye Pan Corresponding author(Shaoxu Li, Ye Pan 通讯作者)

专题命中 效率与蒸馏 :image generation(abstract);分类 cs.CV

AI总结 本文提出了一种高效协调的跨域特征整合框架,通过定制模型和多模型组合提升内容与风格的平衡能力。

Comments https://github.com/lishaoxu1994/DiffStyler, AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08533 2025-12-16 physics.plasm-ph astro-ph.HE 50%

Transition to Petschek Reconnection in Subrelativistic Pair Plasmas: Implications for Particle Acceleration

亚相对论等离子体中向佩斯切克重联转变:对粒子加速的启示

Adam Robbins, Anatoly Spitkovsky

专题命中 效率与蒸馏 :diffusion(abstract)

AI总结 研究亚相对论等离子体中重联几何转变及其对粒子加速的影响,揭示佩斯切克几何与等离子体链在能量谱上的差异。

Comments 13 pages, 9 figures, accepted to ApJ. Comments welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10309 2025-12-16 q-bio.MN cs.LG physics.bio-ph 50%

Tracking large chemical reaction networks and rare events by neural networks

通过神经网络追踪大规模化学反应网络和罕见事件

Jiayu Weng, Xinyi Zhu, Jing Liu, Linyuan Lü, Pan Zhang, Ying Tang

机构 * Institute of Data Science, University of Hong Kong, Hong Kong(数据科学研究所,香港大学) Department of Systems Science, Faculty of Arts and Sciences, Beijing Normal University, Zhuhai 519087, China(艺术与科学学院系统科学系,北京师范大学) School of Physics, University of Electronic Science and Technology of China, Chengdu 611731, China(电子科学与技术大学物理学院) School of Physical Science and Technology, Beijing University of Posts and Telecommunications, Beijing 102206, China(邮电大学物理科学与技术学院) Institute of Theoretical Physics, Chinese Academy of Sciences, Beijing 100190, China(中国科学院理论物理研究所) School of Cyber Science and Technology, University of Science and Technology of China, Hefei 230027, China(科学技术大学网络科学与技术学院) School of Fundamental Physics and Mathematical Sciences, Hangzhou Institute for Advanced Study, UCAS, Hangzhou 310024, China(杭州高等研究院基础物理与数学科学学院) Institute of Fundamental and Frontier Sciences, University of Electronic Science and Technology of China, Chengdu 611731, China(电子科学与技术大学基础与前沿科学研究所) Non-classical Information Science Basic Discipline Research Center of Sichuan Province, University of Electronic Science and Technology of China, Chengdu 611731, China(四川省非经典信息科学基础学科研究中心,电子科学与技术大学)

专题命中 效率与蒸馏 :diffusion(abstract)

AI总结 本文提出通过神经网络高效追踪大规模化学反应网络及罕见事件的方法,结合优化算法和增强采样策略,实现显著加速和高精度建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17756 2025-12-16 cs.LG cs.SY eess.SY 50%

SuperGen: An Efficient Ultra-high-resolution Video Generation System with Sketching and Tiling

SuperGen: 一种高效的超高清视频生成系统,支持草图和分块

Fanjiang Ye, Zepeng Zhao, Yi Mu, Jucheng Shen, Renjie Li, Kaijian Wang, Saurabh Agarwal, Myungjin Lee, Triston Cao, Aditya Akella, Arvind Krishnamurthy, T. S. Eugene Ng, Zhengzhong Tu, Yuke Wang

机构 * Rice University(里士满大学) Carnegie Mellon University(卡内基梅隆大学) University of Illinois Urbana Champaign(伊利诺伊大学厄巴纳-香槟分校) Texas A&M University(德克萨斯农工大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Cisco(思科公司) NVIDIA(英伟达公司) University of Washington(华盛顿大学)

专题命中 效率与蒸馏 :diffusion(abstract)

AI总结 SuperGen通过分块技术实现高效超高清视频生成,无需额外训练,降低计算和内存成本,提升生成速度和质量。

详情

展开后加载摘要…

URL PDF HTML 收藏