arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-07-30 至 2025-07-30 共收录 52 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 34 篇

2507.00090 2025-07-30 cs.LG cs.AI 50%

Generating Heterogeneous Multi-dimensional Data : A Comparative Study

Michael Corbeau, Emmanuelle Claeys, Mathieu Serrurier, Pascale Zaraté

机构 * Institut de Recherche en Informatique de Toulouse(图卢兹信息研究所)

专题命中 扩散模型 :diffusion(abstract)

Comments accepted at IEEE SMC 2025 Vienna

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19517 2025-07-30 cond-mat.soft physics.flu-dyn 50%

Stokes drag on a sphere in a three-dimensional anisotropic porous medium

Andrej Vilfan, Bogdan Cichocki, Jeffrey C. Everts

专题命中 扩散模型 :diffusion(abstract)

Comments 9 pages, 4 figures

Journal ref Phys. Rev. E 112, 015107 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03804 2025-07-30 astro-ph.GA 50%

Metallicity Gradients in Modern Cosmological Simulations I: Tension Between Smooth Stellar Feedback Models and Observations

Alex M. Garcia, Paul Torrey, Aniket Bhagwat, Ruby J. Wright, Qian-hui Chen, Kathryn Grasha, Sophia Ridolfo, Z. S. Hemler, Arnab Sarkar, Priyanka Chakraborty, Erica J. Nelson, Ryan L. Sanders, Tiago Costa, Mark Vogelsberger, Lisa J. Kewley, Sara L. Ellison, Lars Hernquist

专题命中 扩散模型 :diffusion(abstract)

Comments 16 pages, 4 figures, + appendices. Accepted to ApJ

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13343 2025-07-30 cond-mat.mtrl-sci 50%

Design Kinetic Parameters for Improved Resilience of Materials under Irradiation

Mohammadhossein Nahavandian, Eda Aydogan, Jesper Byggmästar, Matheus A. Tunes, Enrique Martinez, Osman El-Atwani

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07196 2025-07-30 quant-ph 50%

Lifetime-Limited and Tunable Emission from Charge-Stabilized Nickel Vacancy Centers in Diamond

I. M. Morris, T. Lühmann, K. Klink, L. Crooks, D. Hardeman, D. J. Twitchen, S. Pezzagna, J. Meijer, S. S. Nicley, J. N. Becker

专题命中 扩散模型 :diffusion(abstract)

Comments 6 pages, 4 figures

Journal ref Phys. Rev. Lett. 135, 043602 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13575 2025-07-30 q-fin.PM 50%

Indifference pricing of pure endowments in a regime-switching market model

Alessandra Cretarola, Benedetta Salterini

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 可控生成 5 篇

2507.21816 2025-07-30 eess.IV 78%

Control Copy-Paste: Controllable Diffusion-Based Augmentation Method for Remote Sensing Few-Shot Object Detection

Yanxing Liu, Jiancheng Pan, Bingchen Zhang

专题命中 可控生成 :diffusion(title,abstract)

Comments 5 Pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08654 2025-07-30 cs.CV 77%

ZeroStereo: Zero-shot Stereo Matching from Single Images

Xianqi Wang, Hao Yang, Gangwei Xu, Junda Cheng, Min Lin, Yong Deng, Jinliang Zang, Yurui Chen, Xin Yang

机构 * Huazhong University of Science and Technology(华中科技大学) Autel Robotics(Autel机器人公司) Optics Valley Laboratory(光学谷实验室)

专题命中 可控生成 :image generation(abstract);diffusion(abstract);inpainting(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21371 2025-07-30 cs.CV 57%

Top2Pano: Learning to Generate Indoor Panoramas from Top-Down View

Zitong Zhang, Suranjan Gautam, Rui Yu

机构 * University of Louisville(路易斯维尔大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments ICCV 2025. Project page: https://top2pano.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21367 2025-07-30 cs.CV 57%

Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic Segmentation

I-Hsiang Chen, Hua-En Chang, Wei-Ting Chen, Jenq-Neng Hwang, Sy-Yen Kuo

机构 * National Taiwan University(台湾大学) University of Washington(华盛顿大学) Microsoft(微软) Chang Gung University(长庚大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20133 2025-07-30 cs.CL cs.AI cs.LG 50%

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering

Anas Mohamed, Azal Ahmad Khan, Xinran Wang, Ahmad Faraz Khan, Shuwen Ge, Saman Bahzad Khan, Ayaan Ahmad, Ali Anwar

机构 * University of Minnesota(明尼苏达大学) Virginia Tech(弗吉尼亚理工大学) Xi’an University of Technology(西安理工大学) Lahore University of Management Sciences(拉合尔管理科学大学) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 可控生成 :text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 图像修复 1 篇

2412.06340 2025-07-30 cs.CV 79%

UniPaint: Unified Space-time Video Inpainting via Mixture-of-Experts

Zhen Wan, Chenyang Qi, Zhiheng Liu, Tao Gui, Yue Ma

机构 * Fudan University(复旦大学) HKUST(香港科技大学) HKU(香港大学)

专题命中 图像修复 :inpainting(title,abstract);分类 cs.CV

Comments ICCV 1st Workshop on Human-Interactive Generation and Editing (poster)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 个性化与一致性 3 篇

2412.03347 2025-07-30 cs.CV cs.AI 77%

DIVE: Taming DINO for Subject-Driven Video Editing

Yi Huang, Wei Xiong, He Zhang, Chaoqi Chen, Jianzhuang Liu, Mingfu Yan, Shifeng Chen

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) vivo AI Lab(vivo AI实验室) NVIDIA Adobe Research(Adobe研究院) Shenzhen University(深圳大学) Southeast University(东南大学) Shenzhen University of Advanced Technology(深圳大学先进技术学院)

专题命中 个性化与一致性 :image generation(abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16418 2025-07-30 cs.CV cs.LG 70%

InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity

Liming Jiang, Qing Yan, Yumin Jia, Zichuan Liu, Hao Kang, Xin Lu

机构 * ByteDance Intelligent Creation Project(字节跳动智能创作项目)

专题命中 个性化与一致性 :image generation(abstract);diffusion(abstract);分类 cs.CV

Comments ICCV 2025 (Highlight). Project page: https://bytedance.github.io/InfiniteYou/ Code and model: https://github.com/bytedance/InfiniteYou

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13568 2025-07-30 cs.CV 57%

LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning

Kaihong Wang, Donghyun Kim, Margrit Betke

机构 * Waymo Boston University(波士顿大学) Korea University(韩国大学)

专题命中 个性化与一致性 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 图像生成评测 5 篇

2411.17489 2025-07-30 cs.CV cs.AI cs.GR cs.LG 62%

Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene Reconstructions

Nicolai Hermann, Jorge Condor, Piotr Didyk

机构 * USI(USI大学) IDSIA(IDSIA研究所)

专题命中 图像生成评测 :inpainting(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21205 2025-07-30 cs.LG cs.AI cs.CV 57%

Learning from Limited and Imperfect Data

Harsh Rangwani

专题命中 图像生成评测 :image generation(abstract);分类 cs.CV

Comments PhD Thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20987 2025-07-30 cs.CV cs.AI 57%

JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1

Xinhan Di, Kristin Qi, Pengqian Yu

机构 * Computer Science, University of Massachusetts Boston(马萨诸塞大学波士顿分校计算机科学系) National University of Singapore(新加坡国立大学)

专题命中 图像生成评测 :diffusion(abstract);分类 cs.CV

Comments WiCV @ ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09081 2025-07-30 cs.CV cs.AI cs.CL 57%

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation

Zheqi He, Yesheng Liu, Jing-shu Zheng, Xuejing Li, Jin-Ge Yao, Bowen Qin, Richeng Xuan, Xi Yang

机构 * BAAI FlagEval Team(BAAI 评测团队)

专题命中 图像生成评测 :text-to-image(abstract);分类 cs.CV

Comments Accepted by ACL 2025 Demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14540 2025-07-30 cs.RO cs.AI cs.CV 57%

IRASim: A Fine-Grained World Model for Robot Manipulation

Fangqi Zhu, Hongtao Wu, Song Guo, Yuxiao Liu, Chilam Cheang, Tao Kong

机构 * Hong Kong University of Science and Technology(香港科技大学) ByteDance Seed(字节跳动种子)

专题命中 图像生成评测 :diffusion(abstract);分类 cs.CV

Comments Opensource, project website: https://gen-irasim.github.io

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 效率与蒸馏 2 篇

2507.21947 2025-07-30 cs.CV cs.AI 57%

Enhancing Generalization in Data-free Quantization via Mixup-class Prompting

Jiwoong Park, Chaeun Lee, Yongseok Choi, Sein Park, Deokki Hong, Jungwook Choi

机构 * HyperAccel KRAFTON SK Telecom Rebellions Hanyang University(翰阳大学)

专题命中 效率与蒸馏 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20650 2025-07-30 cs.CR cs.AI cs.CV 57%

Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model Distribution

Zhicheng Zhang, Peizhuo Lv, Mengke Wan, Jiang Fang, Diandian Guo, Yezeng Chen, Yinlong Liu, Wei Ma, Jiyan Sun, Liru Geng

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of the Chinese Academy of Sciences(中国科学院大学网络安全学院) Nanyang Technological University(南洋理工大学) ShanghaiTech University(上海科技大学)

专题命中 效率与蒸馏 :image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏