arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-09-29 至 2025-09-29 共收录 90 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 57 篇

2507.13019 2025-09-29 cs.RO cs.AI cs.CL cs.CV 57%

Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities

Liuyi Wang, Xinyuan Xia, Hui Zhao, Hanqing Wang, Tai Wang, Yilun Chen, Chengju Liu, Qijun Chen, Jiangmiao Pang

机构 * Tongji University(同济大学) Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) State Key Laboratory of Autonomous Intelligent Unmanned Systems(自主智能无人系统国家重点实验室)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02244 2025-09-29 cs.CV cs.AI 57%

Physics-Guided Motion Loss for Video Generation Model

Bowen Xue, Giuseppe Claudio Guarnera, Shuang Zhao, Zahra Montazeri

机构 * University of Manchester(曼彻斯特大学) University of York(约克大学) University of California, Irvine(加州大学 Irvine 分校)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22604 2025-09-29 quant-ph 50%

Exact solutions of open quantum Brownian motions on the real line for two-level systems

Manuel D. de la Iglesia, Carlos F. Lardizabal

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22098 2025-09-29 physics.optics physics.atom-ph 50%

50 mm $\times$ 50 mm Cesium Atomic Vapor Cell for Terahertz Imaging: Implementation and Application

Bin Zhang, Jun Wan, Tao Li, Xian-Zhe Li, Yu Wu, Qi-Rong Huang, Xin-Yu Yang, Wei Huang, Kai-Qing Zhang, Hai-Xiao Deng

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21717 2025-09-29 astro-ph.SR 50%

MESA Isochrones and Stellar Tracks (MIST) III. The White Dwarf Cooling Sequence

Evan B. Bauer, Aaron Dotter, Charlie Conroy, Tim Cunningham, Minjung Park, Pier-Emmanuel Tremblay

专题命中 扩散模型 :diffusion(abstract)

Comments 13 pages, 10 figures, accepted for publication in ApJS. White dwarf cooling data files available at https://doi.org/10.5281/zenodo.15242046

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21625 2025-09-29 cs.SD cs.AI cs.LG eess.AS 50%

Guiding Audio Editing with Audio Language Model

Zitong Lan, Yiduo Hao, Mingmin Zhao

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21596 2025-09-29 cs.SI physics.soc-ph 50%

Message passing for epidemiological interventions on networks with loops

Erik Weis, Laurent Hébert-Dufresne, Jean-Gabriel Young

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21522 2025-09-29 cs.SD cs.AI 50%

Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training

Naisong Zhou, Saisamarth Rajesh Phaye, Milos Cernak, Tijana Stojkovic, Andy Pearce, Andrea Cavallaro, Andy Harper

机构 * EPFL, Lausanne, CH(瑞士洛桑联邦理工学院) Logitech, Lausanne, CH(瑞士洛桑洛吉科技)

专题命中 扩散模型 :diffusion(abstract)

Comments 5 pages, 2 figures, submitted to ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21457 2025-09-29 cond-mat.soft cond-mat.stat-mech 50%

Viscous Growth Law in Bubble Coarsening: A Molecular Dynamics Perspective

Parameshwaran A, Bhaskar Sen Gupta

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17433 2025-09-29 physics.med-ph 50%

Exploring Machine Learning Models for Physical Dose Calculation in Carbon Ion Therapy Using Heterogeneous Imaging Data -- A Proof of Concept Study

Miriam Schwarze, Hui Khee Looe, Björn Poppe, Pichaya Tappayuthpijarn, Leo Thomas, Hans Rabus

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09498 2025-09-29 cs.AI 50%

SEDM: Scalable Self-Evolving Distributed Memory for Agents

Haoran Xu, Jiacong Hu, Ke Zhang, Lei Yu, Yuxin Tang, Xinyuan Song, Yiqun Duan, Lynn Ai, Bill Shi

机构 * Gradient Zhejiang University(Gradient浙江大学) Gradient South China University of Technology(Gradient南方科技大学) Gradient Waseda University(Gradient武藏大学) Gradient University of Toronto(Gradient多伦多大学) Gradient Rice University(Gradient里士满大学) Gradient Emory University(Gradient埃默里大学) Gradient University of Technology Sydney(Gradient悉尼技术大学)

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17806 2025-09-29 cond-mat.stat-mech cond-mat.dis-nn 50%

Exact results on the hydrodynamics of certain kinetically-constrained hopping processes

Adam J. McRoberts, Vadim Oganesyan, Antonello Scardicchio

专题命中 扩散模型 :diffusion(abstract)

Comments 5 + 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22947 2025-09-29 math.AP 50%

Monotone Multispecies Flows

Lauren Conger, Franca Hoffmann, Eric Mazumdar, Lillian J. Ratliff

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16799 2025-09-29 math.NA cs.NA 50%

Finite element schemes with tangential motion for fourth order geometric curve evolutions in arbitrary codimension

Klaus Deckelnick, Robert Nürnberg

专题命中 扩散模型 :diffusion(abstract)

Comments 30 pages, 10 figures

Journal ref Numer. Math. 157 (2025) 1313--1346

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 可控生成 6 篇

2509.22139 2025-09-29 cs.CV cs.AI 79%

REFINE-CONTROL: A Semi-supervised Distillation Method For Conditional Image Generation

Yicheng Jiang, Jin Yuan, Hua Yuan, Yao Zhang, Yong Rui

机构 * School of Computer Science and Engineering(计算机科学与工程学院) AI Lab(人工智能实验室) Lenovo Research(联想研究院)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments 5 pages,17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21938 2025-09-29 cs.CV cs.AI 70%

SemanticControl: A Training-Free Approach for Handling Loosely Aligned Visual Conditions in ControlNet

Woosung Joung, Daewon Chae, Jinkyu Kim

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Korea University(韩国大学) Electrical and Computer Engineering(电气与计算机工程系) University of Michigan(密歇根大学)

专题命中 可控生成 :text-to-image(abstract);diffusion(abstract);分类 cs.CV

Comments BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14108 2025-09-29 cs.CV 70%

DanceText: A Training-Free Layered Framework for Controllable Multilingual Text Transformation in Images

Zhenyu Yu, Mohd Yamani Idna Idris, Hua Wang, Pei Wang, Rizwan Qureshi, Shaina Raza, Aman Chadha, Yong Xiang, Zhixiang Chen

机构 * Universiti Malaya(马来大学) Zhejiang University of Finance and Economics Dongfang College(浙江财经大学东阳学院) Kunming University of Science and Technology(昆明理工大学) University of Central Florida(佛罗里达中央大学) Toronto metropolitan university(多伦多 Metropolitan 大学) Amazon Web Services(亚马逊网络服务) Deakin University(德肯大学) University of Sheffield(谢菲尔德大学)

专题命中 可控生成 :diffusion(abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21928 2025-09-29 cs.RO cs.AI 67%

SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks

Jialiang Li, Wenzheng Wu, Gaojing Zhang, Yifan Han, Wenzhao Lian

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) ShanghaiTech University(上海科技大学) University of Sussex(Sussex大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 可控生成 :image editing(abstract);inpainting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22169 2025-09-29 cs.CV cs.LG 57%

DragGANSpace: Latent Space Exploration and Control for GANs

Kirsten Odendaal, Neela Kaushik, Spencer Halverson

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 可控生成 :image synthesis(abstract);分类 cs.CV

Comments 6 pages with 7 figures and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22432 2025-09-29 cs.CV 57%

Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy

Yuhao Liu, Tengfei Wang, Fang Liu, Zhenwei Wang, Rynson W. H. Lau

机构 * City University of Hong Kong(香港城市大学) Tencent(腾讯)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted by Siggraph Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 图像修复 1 篇

2509.22570 2025-09-29 cs.AI 75%

UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration

Qi Mao, Tinghan Yang, Jiahao Li, Bin Li, Libiao Jin, Yan Lu

机构 * State Key Laboratory of Media Convergence and Communication(媒体融合与传播国家重点实验室) Communication University of China(中国传媒大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 图像修复 :image generation(abstract);text-to-image(abstract);inpainting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 个性化与一致性 3 篇

2509.21433 2025-09-29 cs.CV cs.AI cs.LG 83%

DyME: Dynamic Multi-Concept Erasure in Diffusion Models with Bi-Level Orthogonal LoRA Adaptation

Jiaqi Liu, Lan Zhang, Xiaoyong Yuan

机构 * Clemson University(克莱姆森大学)

专题命中 个性化与一致性 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03039 2025-09-29 cs.CV cs.AI cs.LG 83%

Leveraging Model Guidance to Extract Training Data from Personalized Diffusion Models

Xiaoyu Wu, Jiaru Zhang, Zhiwei Steven Wu

机构 * Carnegie Mellon University(卡内基梅隆大学) Purdue University(普渡大学)

专题命中 个性化与一致性 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted at the International Conference on Machine Learning (ICML) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21989 2025-09-29 cs.CV 77%

Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation

Abdelrahman Eldesokey, Aleksandar Cvejic, Bernard Ghanem, Peter Wonka

机构 * KAUST, Saudi Arabia(卡士大学)

专题命中 个性化与一致性 :image generation(abstract);diffusion(abstract);image synthesis(abstract);分类 cs.CV

Comments NeurIPS 2025 (Spotlight). Project Page: https://abdo-eldesokey.github.io/mind-the-glitch/

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 图像生成评测 1 篇

2502.02588 2025-09-29 cs.CV 83%

Calibrated Multi-Preference Optimization for Aligning Diffusion Models

Kyungmin Lee, Xiaohang Li, Qifei Wang, Junfeng He, Junjie Ke, Ming-Hsuan Yang, Irfan Essa, Jinwoo Shin, Feng Yang, Yinxiao Li

机构 * Google DeepMind(谷歌DeepMind) KAIST(韩国科学技术院) Google(谷歌) Google Research(谷歌研究) Georgia Institute of Technology(佐治亚理工学院)

专题命中 图像生成评测 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments CVPR 2025, Project page: https://kyungmnlee.github.io/capo.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 效率与蒸馏 5 篇

2506.05096 2025-09-29 cs.CV 79%

Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers

Haosong Liu, Yuge Cheng, Wenxuan Miao, Zihan Liu, Aiyue Chen, Jing Lin, Yiwu Yao, Chen Chen, Jingwen Leng, Yu Feng, Minyi Guo

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Qizhi Institute(上海启智研究院) Huawei Technologies Co.,Ltd(华为技术有限公司)

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21470 2025-09-29 cs.LG cs.AI 78%

Score-based Idempotent Distillation of Diffusion Models

Shehtab Zaman, Chengyan Liu, Kenneth Chiu

专题命中 效率与蒸馏 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15472 2025-09-29 cs.CV 57%

Efficient Multimodal Dataset Distillation via Generative Models

Zhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang, Gaowen Liu, Yan Yan

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Central Florida(中央佛罗里达大学) Cisco Research(思科研究)

专题命中 效率与蒸馏 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09931 2025-09-29 cs.CV cs.AI 57%

Zero-shot detection of buildings in mobile LiDAR using Language Vision Model

June Moh Goo, Zichao Zeng, Jan Boehm

机构 * Department of Civil, Environmental and Geomatic Engineering, University College London(伦敦大学学院建筑、环境与地理工程系)

专题命中 效率与蒸馏 :image generation(abstract);分类 cs.CV

Comments 7 pages, 6 figures, conference

Journal ref The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18578 2025-09-29 cs.CL 50%

Wide-In, Narrow-Out: Revokable Decoding for Efficient and Effective DLLMs

Feng Hong, Geng Yu, Yushi Ye, Haicheng Huang, Huangjie Zheng, Ya Zhang, Yanfeng Wang, Jiangchao Yao

机构 * Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(上海交通大学合作中位网创新中心) Apple(苹果公司) School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)

专题命中 效率与蒸馏 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏