arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-08-04 至 2025-08-04 共收录 49 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 34 篇

2507.22746 2025-08-04 cs.SD cs.CL eess.AS 50%

Next Tokens Denoising for Speech Synthesis

Yanqing Liu, Ruiqing Xue, Chong Zhang, Yufei Liu, Gang Wang, Bohan Li, Yao Qian, Lei He, Shujie Liu, Sheng Zhao

机构 * Microsoft(微软)

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21053 2025-08-04 cs.LG cs.RO 50%

Flow Matching Policy Gradients

David McAllister, Songwei Ge, Brent Yi, Chung Min Kim, Ethan Weber, Hongsuk Choi, Haiwen Feng, Angjoo Kanazawa

机构 * UC Berkeley(伯克利大学) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)

专题命中 扩散模型 :diffusion(abstract)

Comments See our blog post at https://flowreinforce.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20403 2025-08-04 econ.TH cs.LG 50%

A General Framework for Estimating Preferences Using Response Time Data

Federico Echenique, Alireza Fallah, Michael I. Jordan

机构 * Department of Economics, University of California, Berkeley(加州大学伯克利分校经济学系) Department of Computer Science, Rice University(里德大学计算机科学系) Departments of Electrical Engineering and Computer Sciences and Statistics, University of California, Berkeley(加州大学伯克利分校电子工程与计算机科学及统计学系;巴黎Inria) Inria Paris

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18004 2025-08-04 cs.AI 50%

E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI

Yusen Peng, Shuhua Mao

机构 * University of Warwick(沃里克大学) Wuhan University of Technology(武汉理工大学)

专题命中 扩散模型 :diffusion(abstract)

Comments 44 pages,11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03778 2025-08-04 eess.IV q-bio.TO 50%

Generating Novel Brain Morphology by Deforming Learned Templates

Alan Q. Wang, Fangrui Huang, Bailey Trang, Wei Peng, Mohammad Abbasi, Kilian Pohl, Mert Sabuncu, Ehsan Adeli

专题命中 扩散模型 :diffusion(abstract)

Comments Provisional Acceptance at MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01533 2025-08-04 cs.LG cs.AI q-bio.BM 50%

Transformers trained on proteins can learn to attend to Euclidean distance

Isaac Ellmen, Constantin Schneider, Matthew I. J. Raybould, Charlotte M. Deane

专题命中 扩散模型 :diffusion(abstract)

Journal ref Transactions on Machine Learning Research (2025) https://openreview.net/forum?id=mU59bDyqqv

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01300 2025-08-04 physics.plasm-ph astro-ph.SR physics.space-ph 50%

Magnetic Field Amplification and Reconstruction in Rotating Astrophysical Plasmas: Verifying the Roles of $α$ and $β$ in Dynamo Action

Kiwan Park

专题命中 扩散模型 :diffusion(abstract)

Comments 26 pages, 8 figures, submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01355 2025-08-04 cond-mat.stat-mech cond-mat.soft math-ph math.MP math.PR 50%

Spatio-temporal fluctuations in the passive and active Riesz gas on the circle

Leo Touzo, Pierre Le Doussal, Gregory Schehr

专题命中 扩散模型 :diffusion(abstract)

Comments 76 pages, 13 figures

Journal ref J. Stat. Phys. 192, 79 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08242 2025-08-04 math.PR 50%

Exponential ergodicity of stochastic heat equations with Hölder coefficients

Yi Han

专题命中 扩散模型 :diffusion(abstract)

Comments 42 pages. To appear in SIAM Journal on Mathematical Analysis

Journal ref SIAM Journal on Mathematical Analysis Volume 57 Issue 4 August 2025 Pages: 4220 - 4263

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 可控生成 5 篇

2508.00272 2025-08-04 cs.CV 70%

Towards Robust Semantic Correspondence: A Benchmark and Insights

Wenyue Chong

专题命中 可控生成 :diffusion(abstract);image editing(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15867 2025-08-04 cs.CV 70%

PanoLlama: Generating Endless and Coherent Panoramas with Next-Token-Prediction LLMs

Teng Zhou, Xiaoyu Zhang, Yongchuan Tang

机构 * College of Computer Science and Technology(计算机科学与技术学院)

专题命中 可控生成 :image generation(abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00299 2025-08-04 cs.CV cs.AI cs.RO 57%

Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence

Danzhen Fu, Jiagao Hu, Daiguo Zhou, Fei Wang, Zepeng Wang, Wenhua Liao

机构 * MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 可控生成 :inpainting(abstract);分类 cs.CV

Comments ICCV 2025 Workshop (HiGen)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12699 2025-08-04 math.NA cs.NA 50%

Propagation of chaos in infinite horizon and numerical stability for stochastic McKean-Vlasov equations

Zhuoqi Liu, Shuaibin Gao, Chenggui Yuan, Qian Guo

专题命中 可控生成 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12014 2025-08-04 q-fin.RM math.OC q-fin.MF 50%

Singular Control in a Cash Management Model with Ambiguity

Arnon Archankul, Giorgio Ferrari, Tobias Hellmann, Jacco J. J. Thijssen

专题命中 可控生成 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 图像修复 2 篇

2508.00427 2025-08-04 cs.CV cs.AI 83%

Contact-Aware Amodal Completion for Human-Object Interaction via Multi-Regional Inpainting

Seunggeun Chi, Enna Sachdeva, Pin-Hao Huang, Kwonjoon Lee

机构 * Purdue University West Lafayette(普渡大学西拉法叶分校) Honda Research Institute USA(本田美国研究院)

专题命中 图像修复 :inpainting(title,abstract);diffusion(abstract);分类 cs.CV

Comments ICCV 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00418 2025-08-04 cs.CV eess.IV 79%

IN2OUT: Fine-Tuning Video Inpainting Model for Video Outpainting Using Hierarchical Discriminator

Sangwoo Youn, Minji Lee, Nokap Tony Park, Yeonggyoo Jeon, Taeyoung Na

机构 * KAIST(韩国科学技术院) Columbia University(哥伦比亚大学) SK Telecom(SK电信)

专题命中 图像修复 :inpainting(title,abstract);分类 cs.CV

Comments ICIP 2025. Code: https://github.com/sang-w00/IN2OUT

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 个性化与一致性 1 篇

2508.00397 2025-08-04 cs.CV 57%

Video Forgery Detection with Optical Flow Residuals and Spatial-Temporal Consistency

Xi Xue, Kunio Suzuki, Nabarun Goswami, Takuya Shintate

机构 * NABLAS Inc.(NABLAS公司)

专题命中 个性化与一致性 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 图像生成评测 1 篇

2508.00428 2025-08-04 cs.GR cs.HC 57%

Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation

Nan Xiang, Tianyi Liang, Haiwen Huang, Shiqi Jiang, Hao Huang, Yifei Huang, Liangyu Chen, Changbo Wang, Chenhui Li

专题命中 图像生成评测 :text-to-image(abstract);分类 cs.GR

Comments IEEE VIS VAST 2025 ACM 2012 CCS - Human-centered computing, Visualization, Visualization design and evaluation methods

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 其他图像生成 1 篇

2508.00260 2025-08-04 cs.CV cs.MM 81%

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models

Hyundong Jin, Hyung Jin Chang, Eunwoo Kim

机构 * School of Computer Science and Engineering, Chung-Ang University(Chung-Ang 大学计算机科学与工程学院) School of Computer Science, University of Birmingham(布拉德福德大学计算机科学学院)

专题命中 其他图像生成 :generative vision(title,abstract);分类 cs.CV、cs.MM

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏