arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86714 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3486 篇

2504.20996 2025-04-30 cs.CV 57%

X-Fusion: Introducing New Modality to Frozen Large Language Models

Sicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer, Yijun Li, Yuchen Liu, Abhishek Tandon, Eli Shechtman, Krishna Kumar Singh, Yong Jae Lee, Bolei Zhou, Yuheng Li

机构 * University of California, Los Angeles(加州大学洛杉矶分校) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Adobe Research(Adobe研究院)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Project Page: https://sichengmo.github.io/XFusion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10240 2025-04-30 cs.HC cs.AI cs.CV 57%

AltCanvas: A Tile-Based Image Editor with Generative AI for Blind or Visually Impaired People

Seonghee Lee, Maho Kohga, Steve Landau, Sile O'Modhrain, Hari Subramonyam

机构 * Stanford University(斯坦福大学) University of Michigan(密歇根大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07520 2025-04-30 cs.CV 57%

Instruct-ReID: A Multi-purpose Person Re-identification Task with Instructions

Weizhen He, Yiheng Deng, Shixiang Tang, Qihao Chen, Qingsong Xie, Yizhou Wang, Lei Bai, Feng Zhu, Rui Zhao, Wanli Ouyang, Donglian Qi, Yunfeng Yan

机构 * Zhejiang University(浙江大学) Shanghai AI Laboratory(上海人工智能实验室) SenseTime Research(商汤科技研究院) Liaoning Technical University(辽宁工程技术大学) Shanghai Jiao Tong University(上海交通大学) Qing Yuan Research Institute, Shanghai Jiao Tong University(上海交通大学清元研究院)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14684 2025-04-29 eess.IV cs.AI cs.CV 57%

Learning Modality-Aware Representations: Adaptive Group-wise Interaction Network for Multimodal MRI Synthesis

Tao Song, Yicheng Wu, Minhao Hu, Xiangde Luo, Linda Wei, Guotai Wang, Yi Guo, Feng Xu, Shaoting Zhang

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00238 2025-04-18 cs.AI cs.CV cs.LG q-bio.NC 57%

Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem

Declan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata, Kia Ghods, Amogh Joshi, Alexander Ku, Steven M. Frankland, Thomas L. Griffiths, Jonathan D. Cohen, Taylor W. Webb

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11858 2025-04-17 cs.CV 57%

Synthetic Data for Blood Vessel Network Extraction

Joël Mathys, Andreas Plesner, Jorel Elmiger, Roger Wattenhofer

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments Presented at SynthData Workshop at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11455 2025-04-16 cs.CV 57%

SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Junke Wang, Zhi Tian, Xun Wang, Xinyu Zhang, Weilin Huang, Zuxuan Wu, Yu-Gang Jiang

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments technical report, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08586 2025-04-16 cs.CV 57%

PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm

Haoyi Zhu, Honghui Yang, Xiaoyang Wu, Di Huang, Sha Zhang, Xianglong He, Hengshuang Zhao, Chunhua Shen, Yu Qiao, Tong He, Wanli Ouyang

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:2301.00157

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19390 2025-04-15 eess.IV cs.AI cs.CV 57%

Multi-modal Contrastive Learning for Tumor-specific Missing Modality Synthesis

Minjoo Lim, Bogyeong Kang, Tae-Eui Kam

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05508 2025-04-09 cs.CV 57%

PartStickers: Generating Parts of Objects for Rapid Prototyping

Mo Zhou, Josh Myers-Dean, Danna Gurari

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Accepted to CVPR CVEU workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05456 2025-04-09 cs.CV 57%

Generative Adversarial Networks with Limited Data: A Survey and Benchmarking

Omar De Mitri, Ruyu Wang, Marco F. Huber

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05712 2025-04-09 cs.CV cs.AI 57%

MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices

Jianwen Jiang, Gaojie Lin, Zhengkun Rong, Chao Liang, Yongming Zhu, Jiaqi Yang, Tianyun Zhong

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04510 2025-04-08 cs.CV 57%

Attributed Synthetic Data Generation for Zero-shot Domain-specific Image Classification

Shijian Wang, Linxin Song, Ryotaro Shimizu, Masayuki Goto, Hanqian Wu

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04448 2025-04-08 cs.CV eess.IV 57%

Thermoxels: a voxel-based method to generate simulation-ready 3D thermal models

Etienne Chassaing, Florent Forest, Olga Fink, Malcolm Mielle

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03376 2025-04-07 cs.CV 57%

FLAIRBrainSeg: Fine-grained brain segmentation using FLAIR MRI only

Edern Le Bot, Rémi Giraud, Boris Mansencal, Thomas Tourdias, Josè V. Manjon, Pierrick Coupé

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23125 2025-04-01 cs.CV cs.AI 57%

Evaluating Compositional Scene Understanding in Multimodal Generative Models

Shuhao Fu, Andrew Jun Lee, Anna Wang, Ida Momennejad, Trevor Bihl, Hongjing Lu, Taylor W. Webb

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01167 2025-04-01 cs.CV 57%

Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data

Haoxin Li, Boyang Li

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06740 2025-04-01 cs.LG cs.AI cs.CV cs.IR 57%

Sustainable techniques to improve Data Quality for training image-based explanatory models for Recommender Systems

Jorge Paz-Ruza, David Esteban-Martínez, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12202 2025-04-01 cs.AI cs.CV 57%

Nepotistically Trained Generative-AI Models Collapse

Matyas Bohacek, Hany Farid

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Journal ref Published in ICLR DATA-FM Workshop, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20540 2025-03-27 cs.CV 57%

Beyond Intermediate States: Explaining Visual Redundancy through Language

Dingchen Yang, Bowen Cao, Anran Zhang, Weibo Gu, Winston Hu, Guang Chen

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07767 2025-03-25 cs.CV 57%

Learning Visual Generative Priors without Text

Shuailei Ma, Kecheng Zheng, Ying Wei, Wei Wu, Fan Lu, Yifei Zhang, Chen-Wei Xie, Biao Gong, Jiapeng Zhu, Yujun Shen

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Project Page: https://ant-research.github.io/lumos

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03261 2025-03-19 eess.IV cs.CV 57%

Is JPEG AI going to change image forensics?

Edoardo Daniele Cannas, Sara Mandelli, Nataša Popović, Ayman Alkhateeb, Alessandro Gnutti, Paolo Bestagini, Stefano Tubaro

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10779 2025-03-17 cs.CV 57%

The Power of One: A Single Example is All it Takes for Segmentation in VLMs

Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal, James J. Little

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10166 2025-03-14 cs.IR cs.AI cs.MM 57%

ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning

Pengfei Luo, Jingbo Zhou, Tong Xu, Yuan Xia, Linli Xu, Enhong Chen

专题命中 文生图 :text-to-image(abstract);分类 cs.MM

Comments WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08977 2025-03-13 cs.CV 57%

Decoupled Doubly Contrastive Learning for Cross Domain Facial Action Unit Detection

Yong Li, Menglin Liu, Zhen Cui, Yi Ding, Yuan Zong, Wenming Zheng, Shiguang Shan, Cuntai Guan

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing 2025. A novel and elegant feature decoupling method for cross-domain facial action unit detection

Journal ref IEEE Transactions on Image Processing 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08014 2025-03-12 cs.CV cs.AI 57%

MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents

Yun Xing, Nhat Chung, Jie Zhang, Yue Cao, Ivor Tsang, Yang Liu, Lei Ma, Qing Guo

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07491 2025-03-11 eess.IV cs.CV 57%

NeAS: 3D Reconstruction from X-ray Images using Neural Attenuation Surface

Chengrui Zhu, Ryoichi Ishikawa, Masataka Kagesawa, Tomohisa Yuzawa, Toru Watsuji, Takeshi Oishi

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06287 2025-03-11 cs.CV cs.AI 57%

Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding

Seil Kang, Jinyeong Kim, Junhyeok Kim, Seong Jae Hwang

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03651 2025-03-06 cs.CV 57%

DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles

Rui Zhao, Weijia Mao, Mike Zheng Shou

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17710 2025-02-26 cs.AI cs.CL cs.CV cs.LG 57%

Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures

Akhila Yerukola, Saadia Gabriel, Nanyun Peng, Maarten Sap

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments 40 pages, 49 figures

详情

展开后加载摘要…

URL PDF HTML 收藏