arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 4233 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 4233 篇

2310.03602 2025-09-25 cs.CV 57%

Ctrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout Constraints

Chuan Fang, Yuan Dong, Kunming Luo, Xiaotao Hu, Rakesh Shrestha, Ping Tan

机构 * Hong Kong University of Science and Technology(香港理工大学) LightIllusion, China(中国LightIllusion公司) Alibaba Group(阿里巴巴集团) Simon Fraser University, Canada(加拿大Simon Fraser大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17476 2025-09-23 cs.CV 57%

Stable Video-Driven Portraits

Mallikarjun B. R., Fei Yin, Vikram Voleti, Nikita Drobyshev, Maksim Lapin, Aaryaman Vasishta, Varun Jampani

机构 * Stability AI University of Cambridge(剑桥大学) Cantina

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments https://stable-video-driven-portraits.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07570 2025-09-22 cs.CV cs.AI 57%

OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization

Yixuan Yang, Zhen Luo, Tongsheng Ding, Junru Lu, Mingqi Gao, Jinyu Yang, Victor Sanchez, Feng Zheng

机构 * Southern University of Science and Technology(南方科技大学) Shanghai Innovation Institute(上海创新研究院) University of Warwick(沃林顿大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11277 2025-09-19 cs.CV cs.LG 57%

Probing the Representational Power of Sparse Autoencoders in Vision Models

Matthew Lyle Olson, Musashi Hinck, Neale Ratzlaff, Changbai Li, Phillip Howard, Vasudev Lal, Shao-Yen Tseng

机构 * Oracle Intel Labs(英特尔实验室) Oregon State University(俄勒冈州立大学) Thoughtworks

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments ICCV 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20289 2025-09-17 cs.CV 57%

HierRelTriple: Guiding Indoor Layout Generation with Hierarchical Relationship Triplet Losses

Kaifan Sun, Bingchen Yang, Peter Wonka, Jun Xiao, Haiyong Jiang

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11840 2025-09-16 cs.CV 57%

Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation

Tim Lebailly, Vijay Veerabadran, Satwik Kottur, Karl Ridgeway, Michael Louis Iuzzolino

机构 * Meta KU Leuven(鲁汶大学)

专题命中 可控生成 :generative vision(abstract);分类 cs.CV

Comments ICCV 2025 CDEL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19189 2025-09-16 cs.CV cs.LG 57%

An End-to-End Depth-Based Pipeline for Selfie Image Rectification

Ahmed Alhawwary, Janne Mustaniemi, Phong Nguyen-Ha, Janne Heikkilä

机构 * Center for Machine Vision and Signal Analysis (CMVS), University of Oulu(机器视觉与信号分析中心(CMVS)、奥卢大学) Qualcomm AI Research(高通人工智能研究)

专题命中 可控生成 :inpainting(abstract);分类 cs.CV

Comments Accepted at IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02353 2025-09-16 cs.CV cs.AI cs.LG 57%

Semantic Augmentation in Images using Language

Sahiti Yerramilli, Jayant Sravan Tamarapalli, Tanmay Girish Kulkarni, Jonathan Francis, Eric Nyberg

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10466 2025-09-16 cs.CV cs.HC 57%

A Real-Time Diminished Reality Approach to Privacy in MR Collaboration

Christian Fane

专题命中 可控生成 :inpainting(abstract);分类 cs.CV

Comments 50 pages, 12 figures | Demo video: https://youtu.be/udBxj35GEKI?t=499 | Code: https://github.com/c1h1r1i1s1 (multiple repositories)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20724 2025-09-15 cs.CV 57%

Dynamic Motion Blending for Versatile Motion Editing

Nan Jiang, Hongjie Li, Ziye Yuan, Zimo He, Yixin Chen, Tengyu Liu, Yixin Zhu, Siyuan Huang

机构 * Institute for AI, Peking University(人工智能研究院,北京大学) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) Yuanpei College, Peking University(元培学院,北京大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07647 2025-09-10 cs.CV 57%

Semantic Watermarking Reinvented: Enhancing Robustness and Generation Quality with Fourier Integrity

Sung Ju Lee, Nam Ik Cho

机构 * Dept. of ECE & INMC, Seoul National University, Korea(电子工程系及智能纳米材料中心,首尔国立大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted to the IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Project page: https://thomas11809.github.io/SFWMark/ Code: https://github.com/thomas11809/SFWMark

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04819 2025-09-09 eess.IV cs.CV 57%

AURAD: Anatomy-Pathology Unified Radiology Synthesis with Progressive Representations

Shuhan Ding, Jingjing Fu, Yu Gu, Naiteek Sangani, Mu Wei, Paul Vozila, Nan Liu, Jiang Bian, Hoifung Poon

专题命中 可控生成 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04631 2025-09-08 cs.CV 57%

Disentangled Clothed Avatar Generation with Layered Representation

Weitian Zhang, Yichao Yan, Sijing Wu, Manwen Liao, Xiaokang Yang

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能教育部重点实验室、人工智能研究院、上海交通大学) The University of Hong Kong(香港大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments ICCV 2025 highlight, project page: https://olivia23333.github.io/LayerAvatar/

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06002 2025-09-05 cs.CV 57%

Enhanced Generative Data Augmentation for Semantic Segmentation via Stronger Guidance

Quang-Huy Che, Duc-Tri Le, Bich-Nga Pham, Duc-Khai Lam, Vinh-Tiep Nguyen

机构 * University of Information Technology(信息技术大学) Vietnam National University(越南国家大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Published in ICPRAM 2025, ISBN 978-989-758-730-6, ISSN 2184-4313

Journal ref Proceedings of the 14th International Conference on Pattern Recognition Applications and Methods - ICPRAM (2025) 251-262

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01596 2025-09-03 cs.CV cs.AI 57%

O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing

Yuqing Chen, Junjie Wang, Lin Liu, Ruihang Chu, Xiaopeng Zhang, Qi Tian, Yujiu Yang

机构 * Tsinghua University(清华大学) Huawei Inc.(华为公司) Pengcheng National Laboratory(Pengcheng国家实验室)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21090 2025-09-01 cs.CV 57%

Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment

Namu Kim, Wonbin Kweon, Minsoo Kim, Hwanjo Yu

机构 * KT Corporation, South Korea(韩国KT公司) University of Illinois Urbana-Champaign, USA(伊利诺伊大学厄巴纳-香槟分校) Pohang University of Science and Technology (POSTECH), South Korea(浦项科技大学(POSTECH))

专题命中 可控生成 :image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19320 2025-08-29 cs.CV cs.AI 57%

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Ming Chen, Liyuan Cui, Wenyuan Zhang, Haoxian Zhang, Yan Zhou, Xiaohan Li, Songlin Tang, Jiwen Liu, Borui Liao, Hejia Chen, Xiaoqiang Liu, Pengfei Wan

机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队) Zhejiang University(浙江大学) Tsinghua University(清华大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Technical Report. Project Page: https://chenmingthu.github.io/milm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07817 2025-08-29 cs.RO cs.CV 57%

Pixel Motion as Universal Representation for Robot Control

Kanchana Ranasinghe, Xiang Li, E-Ro Nguyen, Cristina Mata, Jongwoo Park, Michael S Ryoo

机构 * Stony Brook University(石溪大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14316 2025-08-29 cs.CV 57%

T-Stars-Poster: A Framework for Product-Centric Advertising Image Design

Hongyu Chen, Min Zhou, Jing Jiang, Jiale Chen, Yang Lu, Zihang Lin, Bo Xiao, Tiezheng Ge, Bo Zheng

机构 * Alibaba Group(阿里巴巴集团) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 可控生成 :image generation(abstract);分类 cs.CV

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19626 2025-08-28 cs.CV 57%

Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model

Jiajun Sun, Zhen Yu, Siyuan Yan, Jason J. Ong, Zongyuan Ge, Lei Zhang

机构 * School of Translational Medicine, Faculty of Medicine, Nursing and Health Sciences, Monash University(转化医学学院,医学、护理与健康科学学院,墨尔本大学) Melbourne Sexual Health Centre, Alfred Health(墨尔本性健康中心,阿尔弗雷德健康中心) AIM for Health Lab, Monash University(健康促进实验室,墨尔本大学) Faculty of IT, Monash University(信息技术学院,墨尔本大学) Faculty of Engineering, Monash University(工程学院,墨尔本大学) Faculty of Infectious and Tropical Diseases, London School of Hygiene and Tropical Medicine(传染病与热带医学学院,伦敦热带医学学院)

专题命中 可控生成 :image synthesis(abstract);分类 cs.CV

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19604 2025-08-28 cs.CV cs.AI 57%

IELDG: Suppressing Domain-Specific Noise with Inverse Evolution Layers for Domain Generalized Semantic Segmentation

Qizhe Fan, Chaoyu Liu, Zhonghua Qiao, Xiaoqin Shen

机构 * School of Sciences, Xi’an University of Technology, Xi’an 710054, China(西安理工大学科学学院) Department of Applied Mathematics(应用数学系) Theoretical Physics, University of Cambridge, UK(理论物理,剑桥大学,英国) Department of Applied Mathematics, The Hong Kong Polytechnic University, Hung Hom, Hong Kong(应用数学系,香港理工大学,九龙)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22141 2025-08-28 cs.CV cs.AI 57%

FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing

Guanwen Feng, Zhiyuan Ma, Yunan Li, Jiahao Yang, Junwei Jing, Qiguang Miao

机构 * Xi’an Key Laboratory of Big Data and Intelligent Vision, Xidian University, Xi’an 710071, China(西安大数据与智能视觉重点实验室,西安电子科技大学,西安710071,中国) Key Laboratory of Collaborative Intelligence Systems, Ministry of Education, Xidian University, Xi’an 710071, China(协同智能系统重点实验室,教育部,西安电子科技大学,西安710071,中国) School of Computer Science and Technology, Xidian University, Xi’an 710071, China(计算机科学与技术学院,西安电子科技大学,西安710071,中国)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19254 2025-08-28 cs.CV cs.AI cs.HC 57%

Real-Time Intuitive AI Drawing System for Collaboration: Enhancing Human Creativity through Formal and Contextual Intent Integration

Jookyung Song, Mookyoung Kang, Nojun Kwak

机构 * SNU(首尔国立大学)

专题命中 可控生成 :image synthesis(abstract);分类 cs.CV

Comments 6 pages, 4 figures, NeurIPS Creative AI Track 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17957 2025-08-27 cs.LG cs.CV 57%

Generative Feature Imputing -- A Technique for Error-resilient Semantic Communication

Jianhao Huang, Qunsong Zeng, Hongyang Du, Kaibin Huang

机构 * Department of Electrical and Electronic Engineering, The University of Hong Kong(电子与电气工程系,香港大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17062 2025-08-26 cs.CV cs.AI 57%

SSG-Dit: A Spatial Signal Guided Framework for Controllable Video Generation

Peng Hu, Yu Gu, Liang Luo, Fuji Ren

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(计算机科学与工程学院,电子科技大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17029 2025-08-26 cs.CV 57%

A Novel Local Focusing Mechanism for Deepfake Detection Generalization

Mingliang Li, Lin Yuanbo Wu, Changhong Liu, Hanxi Li

机构 * Jiangxi Normal University, China(江西师范大学) Swansea University, United Kingdom(斯旺西大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20439 2025-08-26 cs.CV 57%

Image Augmentation Agent for Weakly Supervised Semantic Segmentation

Wangyu Wu, Xianglin Qiu, Siqi Song, Zhenhong Chen, Xiaowei Huang, Fei Ma, Jimin Xiao

机构 * Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学) University of Liverpool(利物浦大学) Microsoft(微软公司)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted at Neurocomputing 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06602 2025-08-26 cs.CL cs.AI cs.LG cs.MM cs.SD eess.AS 57%

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Tianxin Xie, Yan Rong, Pengfei Zhang, Wenwu Wang, Li Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Surrey(萨里大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.MM

Comments The first comprehensive survey on controllable TTS. Accepted to the EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22979 2025-08-26 cs.CV 57%

LumiSculpt: Enabling Consistent Portrait Lighting in Video Generation

Yuxin Zhang, Dandan Zheng, Biao Gong, Shiwen Wang, Jingdong Chen, Ming Yang, Weiming Dong, Changsheng Xu

机构 * MAIS, Institute of Automation, Chinese Academy of Sciecnes(MAIS,自动化研究所,中国科学院) Ant Group(蚂蚁集团)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11183 2025-08-25 cs.CV 57%

OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation

Bohan Li, Xin Jin, Jianan Wang, Yukai Shi, Yasheng Sun, Xiaofeng Wang, Zhuang Ma, Baao Xie, Chao Ma, Xiaokang Yang, Wenjun Zeng

机构 * School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University(电子信息与电气工程学院,上海交通大学) Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波数字孪生研究所,东部技术研究所) PhiGent Robotics, Beijing, China(PhiGent Robotics,北京,中国)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏