arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 4233 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 4233 篇

2411.19324 2024-12-02 cs.CV 57%

Trajectory Attention for Fine-grained Video Motion Control

Zeqi Xiao, Wenqi Ouyang, Yifan Zhou, Shuai Yang, Lei Yang, Jianlou Si, Xingang Pan

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Project Page: xizaoqu.github.io/trajattn/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18000 2024-12-02 cs.CV 57%

Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models

Shuyang Hao, Bryan Hooi, Jun Liu, Kai-Wei Chang, Zi Huang, Yujun Cai

专题命中 可控生成 :image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00233 2024-12-02 cs.SD cs.AI cs.MM eess.AS eess.SP 57%

SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound

Haohe Liu, Xuenan Xu, Yi Yuan, Mengyue Wu, Wenwu Wang, Mark D. Plumbley

专题命中 可控生成 :diffusion(abstract);分类 cs.MM

Comments Accepted by Journal of Selected Topics in Signal Processing (JSTSP). Demo and code: https://haoheliu.github.io/SemantiCodec/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18660 2024-12-02 cs.CV 57%

OOD-HOI: Text-Driven 3D Whole-Body Human-Object Interactions Generation Beyond Training Domains

Yixuan Zhang, Hui Yang, Chuanchen Luo, Junran Peng, Yuxi Wang, Zhaoxiang Zhang

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17949 2024-11-28 cs.CV 57%

ROICtrl: Boosting Instance Control for Visual Generation

Yuchao Gu, Yipin Zhou, Yunfan Ye, Yixin Nie, Licheng Yu, Pingchuan Ma, Kevin Qinghong Lin, Mike Zheng Shou

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Project page at https://roictrl.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17696 2024-11-27 cs.CV 57%

ScribbleLight: Single Image Indoor Relighting with Scribbles

Jun Myeong Choi, Annie Wang, Pieter Peers, Anand Bhattad, Roni Sengupta

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17532 2024-11-27 cs.CV 57%

FTMoMamba: Motion Generation with Frequency and Text State Space Models

Chengjian Li, Xiangbo Shu, Qiongjie Cui, Yazhou Yao, Jinhui Tang

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16327 2024-11-26 cs.CV 57%

CapHDR2IR: Caption-Driven Transfer from Visible Light to Infrared Domain

Jingchao Peng, Thomas Bashford-Rogers, Zhuang Shao, Haitao Zhao, Aru Ranjan Singh, Abhishek Goswami, Kurt Debattista

专题命中 可控生成 :image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15860 2024-11-26 cs.CV 57%

Generalizable Single-view Object Pose Estimation by Two-side Generating and Matching

Yujing Sun, Caiyi Sun, Yuan Liu, Yuexin Ma, Siu Ming Yiu

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted by WACV 2025, not published yet

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14816 2024-11-25 cs.CV cs.RO eess.IV 57%

Unsupervised Multi-view UAV Image Geo-localization via Iterative Rendering

Haoyuan Li, Chang Xu, Wen Yang, Li Mi, Huai Yu, Haijian Zhang

专题命中 可控生成 :image generation(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13836 2024-11-22 cs.CV 57%

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

Lin Sun, Jiale Cao, Jin Xie, Xiaoheng Jiang, Yanwei Pang

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Homepange and code: https://linsun449.github.io/cliper

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18392 2024-11-19 cs.CV 57%

A Reference-Based 3D Semantic-Aware Framework for Accurate Local Facial Attribute Editing

Yu-Kai Huang, Yutong Zheng, Yen-Shuo Su, Anudeepsekhar Bolimera, Han Zhang, Fangyi Chen, Marios Savvides

专题命中 可控生成 :inpainting(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12886 2024-11-19 cs.CV cs.LG 57%

MCM: Multi-condition Motion Synthesis Framework

Zeyu Ling, Bo Han, Yongkang Wongkan, Han Lin, Mohan Kankanhalli, Weidong Geng

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Journal ref International Joint Conference on Artificial Intelligence 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10187 2024-11-18 cs.CV 57%

Try-On-Adapter: A Simple and Flexible Try-On Paradigm

Hanzhong Guo, Jianfeng Zhang, Cheng Zou, Jun Li, Meng Wang, Ruxue Wen, Pingzhong Tang, Jingdong Chen, Ming Yang

专题命中 可控生成 :inpainting(abstract);分类 cs.CV

Comments Image virtual try-on, 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09151 2024-11-15 cs.CV 57%

Mono2Stereo: Monocular Knowledge Transfer for Enhanced Stereo Matching

Yuran Wang, Yingping Liang, Hesong Li, Ying Fu

专题命中 可控生成 :inpainting(abstract);分类 cs.CV

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08328 2024-11-14 cs.CV 57%

Motion Control for Enhanced Complex Action Video Generation

Qiang Zhou, Shaofeng Zhang, Nianzu Yang, Ye Qian, Hao Li

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Project page: https://mvideo-v1.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04598 2024-11-08 cs.CV 57%

Social EgoMesh Estimation

Luca Scofano, Alessio Sampieri, Edoardo De Matteis, Indro Spinelli, Fabio Galasso

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03561 2024-11-07 cs.CV 57%

Estimating Ego-Body Pose from Doubly Sparse Egocentric Video Data

Seunggeun Chi, Pin-Hao Huang, Enna Sachdeva, Hengbo Ma, Karthik Ramani, Kwonjoon Lee

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02229 2024-11-07 cs.CV 57%

FewViewGS: Gaussian Splatting with Few View Matching and Multi-stage Training

Ruihong Yin, Vladimir Yugay, Yue Li, Sezer Karaoglu, Theo Gevers

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted by NeurIPS2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01919 2024-11-05 cs.RO cs.CV 57%

Real-Time Polygonal Semantic Mapping for Humanoid Robot Stair Climbing

Teng Bin, Jianming Yao, Tin Lun Lam, Tianwei Zhang

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted by The 2024 IEEE-RAS International Conference on Humanoid Robots. The code: https://github.com/BTFrontier/polygon_mapping

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01573 2024-11-05 cs.CV cs.LG eess.IV 57%

Conditional Controllable Image Fusion

Bing Cao, Xingxin Xu, Pengfei Zhu, Qilong Wang, Qinghua Hu

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10062 2024-11-05 cs.GR 57%

360° Stereo Image Composition with Depth Adaption

Kun Huang, Fanglue Zhang, Junhong Zhao, Yiheng Li, Neil Dodgson

专题命中 可控生成 :image generation(abstract);分类 cs.GR

Comments A 14 pages paper, this have been puslished at IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS (TVCG) https://ieeexplore.ieee.org/abstract/document/10298822

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12192 2024-11-01 cs.RO cs.AI cs.CV cs.LG 57%

DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control

Zichen Jeff Cui, Hengkai Pan, Aadhithya Iyer, Siddhant Haldar, Lerrel Pinto

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22317 2024-10-30 cs.CV 57%

Multi-Class Textual-Inversion Secretly Yields a Semantic-Agnostic Classifier

Kai Wang, Fei Yang, Bogdan Raducanu, Joost van de Weijer

专题命中 可控生成 :text-to-image(abstract);分类 cs.CV

Comments Accepted in WACV 2025. Code link: https://github.com/wangkai930418/mc_ti

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19079 2024-10-30 cs.CV cs.LG 57%

BIFRÖST: 3D-Aware Image compositing with Language Instructions

Lingxiao Li, Kaixiong Gong, Weihong Li, Xili Dai, Tao Chen, Xiaojun Yuan, Xiangyu Yue

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments NeurIPS 2024, Code Available: https://github.com/lingxiao-li/Bifrost

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06904 2024-10-30 cs.CV 57%

Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks

Muhammad Saif Ullah Khan, Muhammad Ferjad Naeem, Federico Tombari, Luc Van Gool, Didier Stricker, Muhammad Zeshan Afzal

专题命中 可控生成 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05799 2024-10-29 cs.CV 57%

SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-Resolution

Qi Tang, Yao Zhao, Meiqin Liu, Chao Yao

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.02997 2024-10-28 cs.CV 57%

MOGAN: Morphologic-structure-aware Generative Learning from a Single Image

Jinshu Chen, Qihui Xu, Qi Kang, MengChu Zhou

专题命中 可控生成 :image generation(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12647 2024-10-28 cs.CV cs.RO 57%

DiffusionNOCS: Managing Symmetry and Uncertainty in Sim2Real Multi-Modal Category-level Pose Estimation

Takuya Ikeda, Sergey Zakharov, Tianyi Ko, Muhammad Zubair Irshad, Robert Lee, Katherine Liu, Rares Ambrus, Koichi Nishiwaki

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments 8 pages. 9 figures. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.14107 2024-10-25 cs.CV cs.CR 57%

Hijack-GAN: Unintended-Use of Pretrained, Black-Box GANs

Hui-Po Wang, Ning Yu, Mario Fritz

专题命中 可控生成 :image generation(abstract);分类 cs.CV

Comments Accepted by CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏