arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-07-29 至 2025-07-29 共收录 92 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 55 篇

2502.18118 2025-07-29 eess.SP 50%

Generative AI-enabled Wireless Communications for Robust Low-Altitude Economy Networking

Changyuan Zhao, Jiacheng Wang, Ruichen Zhang, Dusit Niyato, Geng Sun, Hongyang Du, Dong In Kim, Abbas Jamalipour

专题命中 扩散模型 :diffusion(abstract)

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05029 2025-07-29 physics.flu-dyn 50%

Assessment of averaged 1D models for column adsorption with 3D computational experiments

Maria Aguareles, Francesc Font

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19799 2025-07-29 cond-mat.mtrl-sci cs.LG 50%

Enhancing Materials Discovery with Valence Constrained Design in Generative Modeling

Mouyang Cheng, Weiliang Luo, Hao Tang, Bowen Yu, Yongqiang Cheng, Weiwei Xie, Ju Li, Heather J. Kulik, Mingda Li

机构 * Quantum Measurement Group, MIT(麻省理工学院量子测量组) Department of Materials Science and Engineering, MIT(麻省理工学院材料科学与工程系) Department of Chemistry, MIT(麻省理工学院化学系) Neutron Scattering Division, Oak Ridge National Laboratory(橡树岭国家实验室中子散射部) Department of Chemistry, Michigan State University(密歇根州立大学化学系) Department of Nuclear Science and Engineering, MIT(麻省理工学院核科学与工程系)

专题命中 扩散模型 :diffusion(abstract)

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 可控生成 13 篇

2507.19939 2025-07-29 cs.CV 88%

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

Jiaze Wang, Rui Chen, Haowang Cui

机构 * Tianjin University(天津大学)

专题命中 可控生成 :text-to-image(title,abstract);diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20855 2025-07-29 cs.CV 57%

Compositional Video Synthesis by Temporal Object-Centric Learning

Adil Kaan Akan, Yucel Yemez

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments 12+21 pages, submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19950 2025-07-29 cs.CV cs.AI 57%

RARE: Refine Any Registration of Pairwise Point Clouds via Zero-Shot Learning

Chengyu Zheng, Jin Huang, Honghua Chen, Mingqiang Wei

机构 * College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) Shenzhen Research Institute, Nanjing University of Aeronautics and Astronautics(南京航空航天大学深圳研究院) School of Data Science, Lingnan University(岭南大学数据科学学院)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21771 2025-07-29 cs.CV 57%

A Unified Image-Dense Annotation Generation Model for Underwater Scenes

Hongkai Lin, Dingkang Liang, Zhenghao Qi, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 可控生成 :text-to-image(abstract);分类 cs.CV

Comments Accepted by CVPR 2025. The code is available at https://github.com/HongkLin/TIDE

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08714 2025-07-29 cs.CV cs.AI 57%

Versatile Multimodal Controls for Expressive Talking Human Animation

Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li, Zixin Zhu, Sanping Zhou, Ming Yang, Le Wang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi'an Jiaotong University, Ant Group(人机混合增强智能国家级实验室,西安交通大学,蚂蚁集团) National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi'an Jiaotong University(人机混合增强智能国家级实验室,西安交通大学) University at Buffalo(布法罗大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted by ACM MM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06438 2025-07-29 cs.CV 57%

Qffusion: Controllable Portrait Video Editing via Quadrant-Grid Attention Learning

Maomao Li, Lijian Lin, Yunfei Liu, Ye Zhu, Yu Li

机构 * International Digital Economy Academy (IDEA)(国际数字经济学院)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments 19 pages

Journal ref TVCG 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19644 2025-07-29 math.OC cs.NA math.NA 50%

Hierarchical clustering and dimensional reduction for optimal control of large-scale agent-based models

Angela Monti, Fasma Diele, Dante Kalise

专题命中 可控生成 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20868 2025-07-29 physics.flu-dyn 50%

Passive control of wing tip vortices through a grooved-tip design

Junchen Tan, Shūji Ōtomo, Ignazio Maria Viola, Yabin Liu

专题命中 可控生成 :diffusion(abstract)

Comments Submitted to Experiments in Fluids

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19910 2025-07-29 eess.SP 50%

Toward Dual-Functional LAWN: Control-Aware System Design for Aerodynamics-Aided UAV Formations

Jun Wu, Weijie Yuan, Qingqing Cheng, Haijia Jin

专题命中 可控生成 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09328 2025-07-29 math.DS 50%

Spatiotemporal SEIQR Epidemic Modeling with Optimal Control for Vaccination, Treatment, and Social Measures

Achraf Zinihi, Matthias Ehrhardt, Moulay Rchid Sidi Ammi

专题命中 可控生成 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09524 2025-07-29 cs.RO 50%

FlowNav: Combining Flow Matching and Depth Priors for Efficient Navigation

Samiran Gode, Abhijeet Nayak, Débora N. P. Oliveira, Michael Krawez, Cordelia Schmid, Wolfram Burgard

机构 * Artificial Intelligence and Robotics Lab, Department of Computer Science and Artificial Intelligence, University of Technology Nuremberg(人工智能与机器人实验室,计算机科学与人工智能系,图腾大学) Inria, Ecole Normale Supérieure, CNRS, PSL Research University(法国国家信息与自动化研究所,巴黎高等师范学校,国家科学研究中心,巴黎综合理工研究学院)

专题命中 可控生成 :diffusion(abstract)

Comments Accepted to IROS'25. Previous version accepted at CoRL 2024 workshop on Learning Effective Abstractions for Planning (LEAP) and workshop on Differentiable Optimization Everywhere: Simulation, Estimation, Learning, and Control

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03655 2025-07-29 cs.LG cs.AI 50%

Geometric Representation Condition Improves Equivariant Molecule Generation

Zian Li, Cai Zhou, Xiyuan Wang, Xingang Peng, Muhan Zhang

机构 * Institute for Artificial Intelligence, Peking University, Beijing, China(北京大学人工智能研究院) School of Intelligence Science and Technology, Peking University, Beijing, China(北京大学智能科学与技术学院) Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MA, USA(麻省理工学院电子工程与计算机科学系) Department of Automation, Tsinghua University, Beijing, China(清华大学自动化系)

专题命中 可控生成 :diffusion(abstract)

Comments Accepted to ICML 2025 as a Spotlight Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09722 2025-07-29 cs.LG cs.SY eess.SY stat.ML 50%

The Pitfalls of Imitation Learning when Actions are Continuous

Max Simchowitz, Daniel Pfrommer, Ali Jadbabaie

机构 * CMU(卡内基梅隆大学) MIT(麻省理工学院)

专题命中 可控生成 :diffusion(abstract)

Comments 98 pages, 2 figures, updated proof sketch

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 图像修复 2 篇

2507.20291 2025-07-29 cs.CV 57%

Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training

Qiaosi Yi, Shuai Li, Rongyuan Wu, Lingchen Sun, Yuhui Wu, Lei Zhang

机构 * The Hong Kong Polytechnic University(香港理工大学) OPPO Research Institute(OPPO研究院)

专题命中 图像修复 :diffusion(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19524 2025-07-29 cs.LG 50%

Kolmogorov Arnold Network Autoencoder in Medicine

Ugo Lomoio, Pierangelo Veltri, Pietro Hiram Guzzi

机构 * DIMES University of Calabria(卡拉布里亚大学DIMES)

专题命中 图像修复 :inpainting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 个性化与一致性 5 篇

2502.21291 2025-07-29 cs.CV 83%

MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing

Xueyun Tian, Wei Li, Bingbing Xu, Yige Yuan, Yuanzhuo Wang, Huawei Shen

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 个性化与一致性 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

Comments This paper have been accepted by ACM MM25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19836 2025-07-29 cs.GR cs.AI cs.CV cs.MM cs.SD 67%

ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion

Xuanchen Wang, Heng Wang, Weidong Cai

机构 * School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)

专题命中 个性化与一致性 :diffusion(abstract);分类 cs.CV、cs.GR、cs.MM

Comments 10 pages, 5 figures, accepted by the 33rd ACM International Conference on Multimedia (ACM MM 2025), demo page: https://choreomuse.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20953 2025-07-29 cs.CV 57%

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation

Dogucan Yaman, Fevziye Irem Eyiokur, Leonard Bärmann, Hazım Kemal Ekenel, Alexander Waibel

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄大学) Istanbul Technical University(伊斯坦布尔技术大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 个性化与一致性 :inpainting(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20721 2025-07-29 cs.CV 57%

AIComposer: Any Style and Content Image Composition via Feature Integration

Haowen Li, Zhenfeng Fan, Zhang Wen, Zhengzhou Zhu, Yunjin Li

专题命中 个性化与一致性 :diffusion(abstract);分类 cs.CV

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19364 2025-07-29 cs.LG cs.AI 50%

CoSTI: Consistency Models for (a faster) Spatio-Temporal Imputation

Javier Solís-García, Belén Vega-Márquez, Juan A. Nepomuceno, Isabel A. Nepomuceno-Chamorro

专题命中 个性化与一致性 :diffusion(abstract)

Comments 14 pages, 7 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 图像生成评测 5 篇

2410.11824 2025-07-29 cs.CV 83%

KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities

Hsin-Ping Huang, Xinyi Wang, Yonatan Bitton, Hagai Taitelbaum, Gaurav Singh Tomar, Ming-Wei Chang, Xuhui Jia, Kelvin C. K. Chan, Hexiang Hu, Yu-Chuan Su, Ming-Hsuan Yang

专题命中 图像生成评测 :image generation(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Project page: https://kitten-project.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03328 2025-07-29 cs.CV cs.AI cs.NE 70%

Visual Enumeration Remains Challenging for Multimodal Generative AI

Alberto Testolin, Kuinan Hou, Marco Zorzi

机构 * Department of General Psychology and Department of Mathematics University of Padova(帕多瓦大学心理学系和数学系) Department of General Psychology University of Padova(帕多瓦大学心理学系) Department of General Psychology and Padova Neuroscience Center University of Padova(帕多瓦大学心理学系和帕多瓦神经科学中心) IRCSS San Camillo Hospital, Venice-Lido(威尼斯利多医院IRCSS桑卡莫医院)

专题命中 图像生成评测 :text-to-image(abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20808 2025-07-29 cs.CV 57%

FantasyID: A dataset for detecting digital manipulations of ID-documents

Pavel Korshunov, Amir Mohammadi, Vidit Vidit, Christophe Ecabert, Sébastien Marcel

专题命中 图像生成评测 :image generation(abstract);分类 cs.CV

Comments Accepted to IJCB 2025; for project page, see https://www.idiap.ch/paper/fantasyid

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21745 2025-07-29 cs.CV 57%

3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models

Yuhan Zhang, Mengchen Zhang, Tong Wu, Tengfei Wang, Gordon Wetzstein, Dahua Lin, Ziwei Liu

机构 * Fudan University(复旦大学) Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Stanford University(斯坦福大学) The Chinese University of Hong Kong(香港中文大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 图像生成评测 :image generation(abstract);分类 cs.CV

Comments Page: https://zyh482.github.io/3DGen-Bench/ ; Code: https://github.com/3DTopia/3DGen-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20880 2025-07-29 cs.SD cs.AI 50%

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment

Renhang Liu, Chia-Yu Hung, Navonil Majumder, Taylor Gautreaux, Amir Ali Bagherzadeh, Chuan Li, Dorien Herremans, Soujanya Poria

专题命中 图像生成评测 :diffusion(abstract)

Comments https://github.com/declare-lab/jamify

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 效率与蒸馏 3 篇

2507.20454 2025-07-29 cs.CV cs.LG 83%

Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis

Zhuokun Chen, Jugang Fan, Zhuowei Yu, Bohan Zhuang, Mingkui Tan

机构 * South China University of Technology(华南理工大学) Pazhou Lab(帕乔实验室) University of California, Davis(加州大学戴维斯分校) Zhejiang University(浙江大学)

专题命中 效率与蒸馏 :image synthesis(title);image generation(abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03515 2025-07-29 cs.CV 79%

Distilling Diffusion Models to Efficient 3D LiDAR Scene Completion

Shengyuan Zhang, An Zhao, Ling Yang, Zejian Li, Chenye Meng, Haoran Xu, Tianrun Chen, AnYang Wei, Perry Pengyun GU, Lingyun Sun

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

Comments This paper is accepted by ICCV'25(Oral), the model and code are publicly available on https://github.com/happyw1nd/ScoreLiDAR

详情

展开后加载摘要…

URL PDF HTML 收藏