arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86714 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3486 篇

1909.04443 2019-09-11 cs.LG stat.ML 67%

Learning Priors for Adversarial Autoencoders

Hui-Po Wang, Wen-Hsiao Peng, Wei-Jan Ko

专题命中 文生图 :text-to-image(abstract);image synthesis(abstract)

Comments Accepted by APSIPA ASC, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.06586 2019-05-17 cs.LG stat.ML 67%

On Conditioning GANs to Hierarchical Ontologies

Hamid Eghbal-zadeh, Lukas Fischer, Thomas Hoch

专题命中 文生图 :image generation(abstract);image synthesis(abstract)

Comments Under review at MLKgraphs2019: http://www.dexa.org/mlkgraphs2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13541 2026-08-14 cs.CV cs.GR 新提交 62%

SCULPT: Subtractive Composition for 3D Part Generation

SCULPT:用于3D部件生成的减法组合方法

Sikuang Li, Chen Yang, Jiemin Fang, Jiazhong Cen, Yuhe Wei, Jichen Pang, Wei Shen, Qi Tian

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei(华为)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 SCULPT是一种3D部件生成框架,通过减法组合方式生成部件,解决了现有方法的边界问题,在PartObjaverse上实现了最优几何性能,还能完成细粒度纹理部件分解。

Comments Project page: https://sculpt-part.github.io/ Code: https://github.com/sculpt-part/SCULPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08831 2026-08-04 cs.GR cs.CV cs.RO 版本更新 62%

DiffPhysCam: Differentiable Physics-Based Camera Simulation for Inverse Rendering and Embodied AI

DiffPhysCam:面向逆渲染与具身智能的可微分基于物理的相机仿真

Bo-Hsun Chen, Nevindu M. Batagoda, Dan Negrut

机构 * Simulation-Based Engineering Lab, University of Wisconsin-Madison(模拟基于工程实验室,威斯康星大学麦迪逊分校)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 本文提出DiffPhysCam,一款可微分基于物理的相机模拟器,解决现有虚拟相机局限,支持正向与逆渲染,经实验验证可提升机器人感知性能,还用于自主地面车辆导航的虚拟实验。

Comments 37 pages, 24 figures, and 5 tables. Code of DiffPhysCam-CamCaliExp: https://github.com/DanielYamChen/DiffPhysCam-CamCaliExp Code of DiffPhysCam-NovelViewSynthesis: https://github.com/DanielYamChen/DiffPhysCam-NovelViewSynthesis Data of DiffPhysCam_Data: https://huggingface.co/datasets/DanielYamChen/DiffPhysCam_Data Simulation video: https://youtu.be/gQwSMrdmHJI

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05290 2026-06-05 cs.CV cs.AI cs.MM 62%

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

模型是否共享安全表示?面向安全视觉生成的跨模型引导

Tobia Poppi, Silvia Cappelletti, Sara Sarto, Florian Schiffers, Garin Kessler, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) University of Pisa(比萨大学) Amazon Prime Video(亚马逊prime视频)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 本文提出首个跨模型安全引导框架,通过源语言模型估计安全方向并迁移至目标生成器,无需目标侧不安全数据即可实现安全控制,且不牺牲生成质量。

Comments Project page: https://aimagelab.github.io/cross-model-safety-representations/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26062 2026-05-26 cs.GR cs.CV 62%

Look Both Ways Before You Cross: Lifting Cross Fields From 2D Visual Priors

过马路前左右看:从2D视觉先验中提取交叉场

Dale Decatur, Jacob Serfaty, Oded Stein, Amir Vaxman, Rana Hanocka

机构 * University of Chicago(芝加哥大学) University of Southern California(南加州大学) University of Edinburgh(爱丁堡大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 提出CrossLift方法,利用文本到图像先验从2D图像中提取方向信号,通过两次平滑插值将其反投影到网格表面,生成语义对齐的交叉场和四边形网格。

Comments Project page at: https://crosslift.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24398 2026-05-26 cs.CV cs.AI cs.GR 62%

VectorArk: Learning Practical Image Vectorization with Rounded Polygon Representation

VectorArk: 学习基于圆角多边形表示的实际图像矢量化

Tarun Gehlaut, Difan Liu, Charu Bansal, Krutik Malani, Souymodip Chakraborty, Ankit Phogat, Matthew Fisher, Vineet Batra

机构 * Adobe

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 提出VectorArk模型,采用圆角多边形表示和退化模型,实现鲁棒且实用的图像矢量化,在多个数据集上取得优越的几何完整性和伪影抑制效果。

Comments CVPR 2026. Project page: https://vectorark.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03558 2026-04-30 cs.CV cs.AI cs.MM 62%

ELIQ: A Label-Free Framework for Quality Assessment of Evolving AI-Generated Images

ELIQ:一种无需标签的进化AI生成图像质量评估框架

Xinyue Li, Zhiming Xu, Min Tang, Zhaolin Cai, Sijing Wu, Xiongkuo Min, Yitong Chen, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Xi'an Jiaotong University(西安交通大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 ELIQ通过自动构建正负样本对,利用预训练多模态模型提升质量评估,实现无需人工标注的高质量评估,优于现有无标签方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22783 2026-04-07 cs.IR cs.CV cs.LG cs.MM cs.SD 62%

Compact Hypercube Embeddings for Fast Text-based Wildlife Observation Retrieval

紧凑超立方体嵌入用于快速基于文本的野生动物观察检索

Ilyass Moummad, Marius Miron, David Robinson, Kawtar Zaher, Hervé Goëau, Olivier Pietquin, Pierre Bonnet, Emmanuel Chemla, Matthieu Geist, Alexis Joly

机构 * Inria, LIRMM, UM(法国国家信息与自动化研究所,蒙彼利埃计算机科学实验室,蒙彼利埃大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 本文提出紧凑超立方体嵌入框架,通过二进制表示实现高效文本检索,提升大规模野生动物图像和音频数据库的检索效率与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12459 2026-03-17 cs.CV cs.GR 62%

From Particles to Fields: Reframing Photon Mapping with Continuous Gaussian Photon Fields

从粒子到场:用连续高斯光场重新框架光子映射

Jiachen Tao, Benjamin Planche, Van Nguyen Nguyen, Junyi Wu, Yuchun Liu, Haoxuan Wang, Zhongpai Gao, Gengyu Zhang, Meng Zheng, Feiran Wang, Anwesa Choudhuri, Zhenghao Zhao, Weitai Kang, Terrence Chen, Yan Yan, Ziyan Wu

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 本文提出高斯光场,通过学习表示将光子分布编码为各向异性3D高斯体,提升多视角渲染效率,实现光子级精度与神经场景表示的结合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17623 2026-03-04 cs.MM cs.CV 62%

Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?

合成感知:生成图像能否解锁潜在的视觉先验以用于以文本为中心的推理?

Yuesheng Huang, Peng Zhang, Xiaoxin Wu, Riliang Liu, Jiaqi Liang

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 本文探讨生成图像能否解锁潜在视觉先验以提升文本中心推理,通过多模态融合架构和提示工程策略实现性能提升。

Comments Accepted as a poster at the International Conference on Machine Learning (ICML 2025) NewInML Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19750 2026-01-29 cs.MM cs.CV cs.IR 62%

Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues

在电子商务目录中评估多模态大语言模型用于缺失模态补全

Junchen Fu, Wenhao Deng, Kaiwen Zheng, Ioannis Arapakis, Yu Ye, Yongxin Ni, Joemon M. Jose, Xuri Ge

机构 * University of Glasgow(格拉斯哥大学) Telefónica Scientific Research(电信科研机构) National University of Singapore(新加坡国立大学) Shandong University(山东大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 本文研究了多模态大语言模型在电子商务目录中补全缺失模态的能力,通过MMPCBench基准测试发现MLLMs在细粒度对齐上存在不足,并探索了GRPO方法以提升补全效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03052 2025-12-04 cs.GR cs.CV 62%

LATTICE: Democratize High-Fidelity 3D Generation at Scale

LATTICE:大规模实现高保真3D生成

Zeqiang Lai, Yunfei Zhao, Zibo Zhao, Haolin Liu, Qingxiang Lin, Jingwei Huang, Chunchao Guo, Xiangyu Yue

机构 * MMLab, CUHK(CUHK人工智能实验室) Tencent Hunyuan(腾讯文言)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 LATTICE通过VoxSet半结构化表示实现高效高保真3D生成,支持任意分辨率解码和灵活推理,达到最先进的性能。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05229 2025-10-29 cs.CV cs.MM 62%

Does CLIP perceive art the same way we do?

Andrea Asperti, Leonardo Dessì, Maria Chiara Tonetti, Nico Wu

机构 * Dept. of Informatics (DISI) University of Bologna(信息学院(DISI)博洛尼亚大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.MM

Journal ref Proceedings of IEEE International Conference on Content-Based Multimedia Indexing (IEEE CBMI 2025), Dublin, Ireland, 22-24 October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26055 2025-10-01 cs.GR cs.CV cs.LG 62%

GaussEdit: Adaptive 3D Scene Editing with Text and Image Prompts

Zhenyu Shu, Junlong Yu, Kai Chao, Shiqing Xin, Ligang Liu

机构 * School of Computer and Data Engineering, NingboTech University(计算机与数据工程学院,宁波科技学院) Ningbo Institute, Zhejiang University(浙江大学宁波学院) School of Software Technology, Zhejiang University(软件技术学院,浙江大学) School of Big Data and Artificial Intelligence Management, Xi’an Jiaotong University(大数据与人工智能管理学院,西安交通大学) School of Computer Science and Technology, ShanDong University(计算机科学与技术学院,山东大学) Graphics & Geometric Computing Laboratory, School of Mathematical Sciences, University of Science and Technology of China(图形与几何计算实验室,数学科学学院,中国科学技术大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Journal ref IEEE Transactions on Visualization and Computer Graphics. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15773 2025-08-22 cs.CV cs.GR cs.LG 62%

Scaling Group Inference for Diverse and High-Quality Generation

Gaurav Parmar, Or Patashnik, Daniil Ostashev, Kuan-Chieh Wang, Kfir Aberman, Srinivasa Narasimhan, Jun-Yan Zhu

机构 * Carnegie Mellon University(卡内基梅隆大学) Snap Research(Snap研究公司) Tel Aviv University(特拉维夫大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

Comments Project website: https://www.cs.cmu.edu/~group-inference, GitHub: https://github.com/GaParmar/group-inference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14624 2025-07-22 cs.GR cs.CV 62%

Real-Time Scene Reconstruction using Light Field Probes

Yaru Liu, Derek Nowrouzezahri, Morgan Mcguire

机构 * University of Cambridge(剑桥大学) McGill University(麦吉尔大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20875 2025-06-27 cs.GR cs.CV 62%

3DGH: 3D Head Generation with Composable Hair and Face

Chengan He, Junxuan Li, Tobias Kirschstein, Artem Sevastopolsky, Shunsuke Saito, Qingyang Tan, Javier Romero, Chen Cao, Holly Rushmeier, Giljoo Nam

机构 * Yale University(耶鲁大学) Meta Codec Avatars Lab(Meta 编码人脸实验室) Technical University of Munich(慕尼黑技术大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Accepted to SIGGRAPH 2025. Project page: https://c-he.github.io/projects/3dgh/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15891 2025-05-20 cs.GR cs.CV 62%

TexPro: Text-guided PBR Texturing with Procedural Material Modeling

Ziqiang Dang, Wenqi Dong, Zesong Yang, Bangbang Yang, Liang Li, Yuewen Ma, Zhaopeng Cui

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

Comments Accepted by CVM 2025 and CVMJ (Computational Visual Media Journal)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19296 2025-03-26 cs.CV cs.MM 62%

Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval

Haoqiang Lin, Haokun Wen, Xuemeng Song, Meng Liu, Yupeng Hu, Liqiang Nie

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18588 2025-01-31 cs.HC cs.AI cs.CV cs.MM 62%

Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching

David Chuan-En Lin, Hyeonsu B. Kang, Nikolas Martelaro, Aniket Kittur, Yan-Ying Chen, Matthew K. Hong

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

Comments Accepted to CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16096 2024-11-26 cs.CV cs.AI cs.MM 62%

ENCLIP: Ensembling and Clustering-Based Contrastive Language-Image Pretraining for Fashion Multimodal Search with Limited Data and Low-Quality Images

Prithviraj Purushottam Naik, Rohit Agarwal

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16254 2024-07-25 cs.CV cs.AI cs.CL cs.MM 62%

Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models

Samuele Poppi, Tobia Poppi, Federico Cocchi, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15155 2024-07-23 cs.CV cs.AI cs.MM 62%

Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification

Yunyi Xuan, Weijie Chen, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACMMM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02430 2024-07-03 cs.CV cs.AI cs.GR cs.LG 62%

Meta 3D TextureGen: Fast and Consistent Texture Generation for 3D Objects

Raphael Bensadoun, Yanir Kleiman, Idan Azuri, Omri Harosh, Andrea Vedaldi, Natalia Neverova, Oran Gafni

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16759 2024-03-21 cs.CV cs.GR 62%

StyleHumanCLIP: Text-guided Garment Manipulation for StyleGAN-Human

Takato Yoshikawa, Yuki Endo, Yoshihiro Kanamori

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments VISIAPP 2024, project page: https://www.cgg.cs.tsukuba.ac.jp/~yoshikawa/pub/style_human_clip/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09249 2023-12-15 cs.CV cs.GR 62%

ZeroRF: Fast Sparse View 360° Reconstruction with Zero Pretraining

Ruoxi Shi, Xinyue Wei, Cheng Wang, Hao Su

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Project page: https://sarahweiii.github.io/zerorf/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18491 2023-12-01 cs.CV cs.AI cs.GR cs.LG 62%

ZeST-NeRF: Using temporal aggregation for Zero-Shot Temporal NeRFs

Violeta Menéndez González, Andrew Gilbert, Graeme Phillipson, Stephen Jolly, Simon Hadfield

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments VUA BMVC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03335 2023-11-07 cs.CV cs.GR 62%

Cross-Image Attention for Zero-Shot Appearance Transfer

Yuval Alaluf, Daniel Garibi, Or Patashnik, Hadar Averbuch-Elor, Daniel Cohen-Or

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

Comments Project page: https://garibida.github.io/cross-image-attention

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12302 2023-09-22 cs.CV cs.GR 62%

Text-Guided Vector Graphics Customization

Peiying Zhang, Nanxuan Zhao, Jing Liao

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

Comments Accepted by SIGGRAPH Asia 2023. Project page: https://intchous.github.io/SVGCustomization

详情

展开后加载摘要…

URL PDF HTML 收藏