arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-08-13 至 2025-08-13 共收录 44 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 31 篇

2508.08387 2025-08-13 math.DS 50%

Wave Propagation Dynamics via Lattice Difference Equations

Eddy Kwessi

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08330 2025-08-13 math.DS cs.SY eess.SY physics.class-ph 50%

On Irreversibility and Stochastic Systems: Part One

Giorgio Picci

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04930 2025-08-13 physics.flu-dyn physics.comp-ph 50%

Simulation of Non-Premixed, Supersonic Combustion using the Discontinuous Galerkin Method on Fully Unstructured Grids

Cal J. Rising, Eric J. Ching, Ryan F. Johnson

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 可控生成 4 篇

2508.08949 2025-08-13 cs.CV 79%

Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation

Ao Ma, Jiasong Feng, Ke Cao, Jing Wang, Yun Wang, Quanwei Zhang, Zhanjie Zhang

机构 * JD.com, Inc.(京东公司)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10848 2025-08-13 cs.CV cs.AI cs.LG cs.MM cs.SD eess.AS 62%

3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control

Xuanmeng Sha, Liyun Zhang, Tomohiro Mashita, Naoya Chiba, Yuki Uranishi

机构 * The University of Osaka(大阪大学) Osaka Electro-Communication University(大阪电通信大学) Osaka University(大阪大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08588 2025-08-13 cs.CV eess.IV 57%

RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space

Jingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao, Lei Sun, Yichen Qian, Weihua Chen, Fan Wang

机构 * DAMO Academy, Alibaba Group(阿里达摩院) Hupan Lab(华盘实验室) INSAIT Zhejiang University(浙江大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Project page: https://jingyunliang.github.io/RealisMotion

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08269 2025-08-13 cs.RO cs.AI 50%

emg2tendon: From sEMG Signals to Tendon Control in Musculoskeletal Hands

Sagar Verma

专题命中 可控生成 :diffusion(abstract)

Comments Accepted in Robotics: Science and Systems (RSS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 图像修复 2 篇

2508.08890 2025-08-13 eess.AS cs.SD 88%

Transient Noise Removal via Diffusion-based Speech Inpainting

Mordehay Moradi, Sharon Gannot

机构 * Faculty of Engineering, Bar-Ilan University(巴伊兰大学工程学院)

专题命中 图像修复 :inpainting(title,abstract);diffusion(title,abstract)

Comments 23 pages, 3 figures, signal processing paper on speech inpainting

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08953 2025-08-13 eess.AS cs.SD 50%

Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation

Soo-Whan Chung, Min-Seok Choi

机构 * NAVER Cloud(NAVER云)

专题命中 图像修复 :diffusion(abstract)

Comments Accepted to INTERSPEECH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 个性化与一致性 2 篇

2508.08812 2025-08-13 cs.CV 85%

TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models

Yuqi Peng, Lingtao Zheng, Yufeng Yang, Yi Huang, Mingfu Yan, Jianzhuang Liu, Shifeng Chen

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院) Northeastern University(东北大学) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 个性化与一致性 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09045 2025-08-13 cs.CV 57%

Per-Query Visual Concept Learning

Ori Malca, Dvir Samuel, Gal Chechik

机构 * Bar‑Ilan University(巴伊兰大学) OriginAI NVIDIA

专题命中 个性化与一致性 :text-to-image(abstract);分类 cs.CV

Comments Project page is at https://per-query-visual-concept-learning.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 图像生成评测 1 篇

2311.18100 2025-08-13 cond-mat.soft physics.bio-ph physics.data-an q-bio.QM 50%

Quantitative evaluation of methods to analyze motion changes in single-particle experiments

Gorka Muñoz-Gil, Harshith Bachimanchi, Jesús Pineda, Benjamin Midtvedt, Gabriel Fernández-Fernández, Borja Requena, Yusef Ahsini, Solomon Asghar, Jaeyong Bae, Francisco J. Barrantes, Steen W. B. Bender, Clément Cabriel, J. Alberto Conejero, Marc Escoto, Xiaochen Feng, Rasched Haidari, Nikos S. Hatzakis, Zihan Huang, Ignacio Izeddin, Hawoong Jeong, Yuan Jiang, Jacob Kæstel-Hansen, Judith Miné-Hattab, Ran Ni, Junwoo Park, Xiang Qu, Lucas A. Saavedra, Hao Sha, Nataliya Sokolovska, Yongbing Zhang, Giorgio Volpe, Maciej Lewenstein, Ralf Metzler, Diego Krapf, Giovanni Volpe, Carlo Manzo

专题命中 图像生成评测 :diffusion(abstract)

Comments 37 pages, 8 figures. This is the author's version of the article published in Nature Communications under CC BY 4.0. The final published version is available at https://doi.org/10.1038/s41467-025-61949-x

Journal ref Nat Commun 16, 6749 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 效率与蒸馏 2 篇

2508.08978 2025-08-13 cs.CV 57%

TaoCache: Structure-Maintained Video Generation Acceleration

Zhentao Fan, Zongzuo Wang, Weiwei Zhang

机构 * Huawei Inc.(华为公司)

专题命中 效率与蒸馏 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10021 2025-08-13 cs.CV 57%

Masked Autoencoder Self Pre-Training for Defect Detection in Microelectronics

Nikolai Röhrich, Alwin Hoffmann, Richard Nordsieck, Emilio Zarbali, Alireza Javanmardi

机构 * XITASO GmbH, Germany(德国XITASO GmbH) Institute of Informatics, LMU Munich, Germany(德国慕尼黑大学信息学院) Munich Center for Machine Learning (MCML), Germany(德国慕尼黑机器学习中心)

专题命中 效率与蒸馏 :image generation(abstract);分类 cs.CV

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏