arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-11-27 至 2025-11-27 共收录 76 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 8 篇

2502.02465 2025-11-27 cs.CV 79%

Towards Consistent and Controllable Image Synthesis for Face Editing

面向面部编辑的图像合成一致性与可控性研究

Mengting Wei, Tuomas Varanka, Yante Li, Xingxun Jiang, Huai-Qian Khor, Guoying Zhao

机构 * Center for Machine Vision and Signal Analysis, Faculty of Information Technology and Electrical Engineering, University of Oulu(机器视觉与信号分析中心,信息科技与电气工程学院,奥卢大学) Key Laboratory of Child Development and Learning Science of Ministry of Education, School of Biological Sciences and Medical Engineering, Southeast University(教育部儿童发展与学习科学重点实验室,生物科学与医学工程学院,东南大学)

专题命中 可控生成 :image synthesis(title);diffusion(abstract);分类 cs.CV

AI总结 RigFace通过结合SD模型和3D面部模型,实现面部图像的光照、表情和姿态可控,提升身份保持和图像真实度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21024 2025-11-27 cs.CV 70%

CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching

CameraMaster: 用于摄影修饰的统一相机语义-参数控制

Qirui Yang, Yang Yang, Ying Zeng, Xiaobin Hu, Bo Li, Huanjing Yue, Jingyu Yang, Peng-Tao Jiang

机构 * Tianjin University(天津大学) vivo Mobile Communication Co., Ltd(vivo移动通信有限公司) NUS(国立大学)

专题命中 可控生成 :diffusion(abstract);image editing(abstract);分类 cs.CV

AI总结 CameraMaster通过统一相机感知框架实现精确摄影修饰,通过解耦相机指令和参数嵌入,提升多参数控制与语义-参数对齐能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21146 2025-11-27 cs.MM cs.CV cs.SD 62%

AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control

AV-Edit: 通过音频-视觉语义联合控制实现多模态生成式声音效果编辑

Xinyue Guo, Xiaoran Yang, Lipan Zhang, Jianxuan Yang, Zhao Wang, Jian Luan

专题命中 可控生成 :diffusion(abstract);分类 cs.CV、cs.MM

AI总结 AV-Edit通过联合利用视觉、音频和文本语义,实现多模态生成式声音效果编辑,提升音频编辑的精度和质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21043 2025-11-27 cs.CV 57%

PG-ControlNet: A Physics-Guided ControlNet for Generative Spatially Varying Image Deblurring

PG-ControlNet:一种用于生成空间变化图像去模糊的物理引导ControlNet

Hakki Motorcu, Mujdat Cetin

机构 * 1 Computer Science Department, University of Rochester 2 Goergen Institute for Data Science \& AIS, University of Rochester Computer Engineering Department, University of Rochester

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

AI总结 PG-ControlNet通过结合物理约束与生成模型,提升空间变化图像去模糊的物理准确性和感知真实感。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20784 2025-11-27 cs.CV 57%

One Patch is All You Need: Joint Surface Material Reconstruction and Classification from Minimal Visual Cues

一个补丁就够了:从最小的视觉线索联合表面材料重建与分类

Sindhuja Penchala, Gavin Money, Gabriel Marques, Samuel Wood, Jessica Kirschman, Travis Atkison, Shahram Rahimi, Noorbakhsh Amiri Golilarz

机构 * Department of Computer Science, The University of Alabama(计算机科学系,阿拉巴马大学)

专题命中 可控生成 :inpainting(abstract);分类 cs.CV

AI总结 SMARC通过单个补丁实现表面材料重建与分类,以应对稀疏视觉输入下的挑战,取得最佳性能。

Comments 9 pages,3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21469 2025-11-27 math.AP 50%

A Hamilton-Jacobi Framework in a Field-Road System with Unidirectional Advection under Wentzell-Type Boundary Condition

一个带有单向输运和Wentzell型边界条件的场-路系统中的Hamilton-Jacobi框架

Xinye Xiao, Haomin Huang

专题命中 可控生成 :diffusion(abstract)

AI总结 本文提出了一种用于分析场-路系统中单向输运和Wentzell型边界条件的Hamilton-Jacobi框架,通过变分方法和数值模拟揭示了传播行为的临界转变。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 图像修复 1 篇

2511.20152 2025-11-27 cs.CV 77%

Restora-Flow: Mask-Guided Image Restoration with Flow Matching

Restora-Flow:基于掩码的图像恢复与流匹配

Arnela Hadzic, Franz Thaler, Lea Bogensperger, Simon Johannes Joham, Martin Urschler

专题命中 图像修复 :image generation(abstract);diffusion(abstract);inpainting(abstract);分类 cs.CV

AI总结 Restora-Flow通过降质掩码引导流匹配采样并结合轨迹校正机制,实现高效且高质量的图像恢复。

Comments Accepted for WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 个性化与一致性 3 篇

2502.18477 2025-11-27 cs.IR cs.AI cs.LG 82%

Personalized Image Generation for Recommendations Beyond Catalogs

推荐超越目录的个性化图像生成

Gabriel Patron, Zhiwei Xu, Ishan Kapnadak, Felipe Maia Polo

机构 * University of Michigan(密歇根大学)

专题命中 个性化与一致性 :image generation(title,abstract);diffusion(abstract)

AI总结 REBECA通过直接学习隐含反馈信号实现高效个性化图像生成,无需微调扩散模型,提升大规模用户群体的个性化效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15711 2025-11-27 cs.AI q-bio.NC 67%

Semi-supervised Multimodal Representation Learning through a Global Workspace

通过全局工作空间实现半监督多模态表示学习

Benjamin Devillers, Léopold Maytié, Rufin VanRullen

机构 * CerCo, CNRS UMR 5549, Université de Toulouse and ANITI, Artificial and Natural Intelligence Toulouse Institute(CerCo、CNRS UMR 5549、图卢兹大学和ANITI人工智能与自然智能图卢兹研究所)

专题命中 个性化与一致性 :image generation(abstract);text-to-image(abstract)

AI总结 本文提出了一种受全局工作空间概念启发的神经网络架构,通过自监督学习实现多模态表示对齐与转换,显著减少对匹配数据的需求。

Comments Under review

Journal ref IEEE Transactions on Neural Networks and Learning Systems 36 (5), 7843-7857 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21216 2025-11-27 cs.CR 50%

AuthenLoRA: Entangling Stylization with Imperceptible Watermarks for Copyright-Secure LoRA Adapters

AuthenLoRA: 将隐秘水印与风格化结合以实现版权安全的LoRA适配器

Fangming Shi, Li Li, Kejiang Chen, Guorui Feng, Xinpeng Zhang

专题命中 个性化与一致性 :diffusion(abstract)

AI总结 AuthenLoRA通过在LoRA训练过程中嵌入隐秘水印,实现版权安全的风格化生成,提升水印传播的鲁棒性和降低假阳性率。

Comments 16 pages, 7 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 图像生成评测 1 篇

2511.20937 2025-11-27 cs.AI cs.CL cs.CV cs.RO 57%

ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction

ENACT:通过视角交互的世界建模评估具身认知

Qineng Wang, Wenlong Huang, Yu Zhou, Hang Yin, Tianwei Bao, Jianwen Lyu, Weiyu Liu, Ruohan Zhang, Jiajun Wu, Li Fei-Fei, Manling Li

机构 * Northwestern University(西北大学) Stanford University(斯坦福大学) UCLA(加州大学洛杉矶分校)

专题命中 图像生成评测 :image synthesis(abstract);分类 cs.CV

AI总结 ENACT通过视角交互的世界建模评估具身认知,揭示VLMs在长周期交互中与人类表现差距扩大,且模型在逆向任务中表现更优。

Comments Preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 效率与蒸馏 4 篇

2505.16687 2025-11-27 cs.CV eess.IV 79%

One-Step Diffusion-Based Image Compression with Semantic Distillation

一步扩散式图像压缩与语义蒸馏

Naifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li, Yuan Zhang, Yan Lu

机构 * Communication University of China(中国通信大学) University of Science and Technology of China(中国科学技术大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

AI总结 OneDC通过一步扩散生成和语义蒸馏机制实现高效图像压缩,达到SOTA感知质量,比特率降低39%且解码速度提升20倍。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20705 2025-11-27 cs.LG cs.AI stat.ML 78%

Solving Diffusion Inverse Problems with Restart Posterior Sampling

利用重启后验抽样解决扩散逆问题

Bilal Ahmed, Joseph G. Makin

专题命中 效率与蒸馏 :diffusion(title,abstract)

AI总结 RePS通过重启后验抽样高效解决线性和非线性逆问题,避免反向传播,提升收敛速度和重建质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21475 2025-11-27 cs.CV 57%

MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices

MobileI2V: 一种适用于移动设备的快速高分辨率图像到视频生成方法

Shuai Zhang, Bao Tang, Siyuan Yu, Yueting Zhu, Jingfeng Yao, Ya Zou, Shanglin Yuan, Li Yu, Wenyu Liu, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 效率与蒸馏 :diffusion(abstract);分类 cs.CV

AI总结 MobileI2V通过轻量级扩散模型和优化策略,在移动设备上实现快速高分辨率图像到视频生成,生成速度提升10倍,质量与现有模型相当。

Comments Our Demo and code:https://github.com/hustvl/MobileI2V

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21136 2025-11-27 cs.CV cs.AI 57%

Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning

高效训练人类视频生成的熵引导优先级渐进学习

Changlin Li, Jiawei Zhang, Shuhao Liu, Sihao Lin, Zeyi Shi, Zhihui Li, Xiaojun Chang

机构 * Stanford University(斯坦福大学) North China Electric Power University(华北电力大学) University of Adelaide(阿德莱德大学) University of Technology Sydney(悉尼技术大学) University of Science and Technology of China(中国科学技术大学)

专题命中 效率与蒸馏 :diffusion(abstract);分类 cs.CV

AI总结 本文提出熵引导优先级渐进学习方法,通过条件熵膨胀和自适应渐进计划,高效训练扩散模型生成人类视频,实现训练速度提升和内存消耗降低。

Comments Project page: https://github.com/changlin31/Ent-Prog

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 其他图像生成 1 篇

2511.21057 2025-11-27 cs.CV 79%

Long-Term Alzheimers Disease Prediction: A Novel Image Generation Method Using Temporal Parameter Estimation with Normal Inverse Gamma Distribution on Uneven Time Series

长期阿尔茨海默病预测:一种利用正常逆伽玛分布进行时间参数估计的新型图像生成方法

Xin Hong, Xinze Sun, Yinhao Li, Yen-Wei Chen

机构 * College of Computer Science and Technology, Huaqiao University(华侨大学计算机科学与技术学院) Key Laboratory of Computer Vision and Machine Learning in Fujian Province(福建省计算机视觉与机器学习重点实验室) College of Information Science and Engineering,Ritsumeikan University(立命馆大学信息科学与工程学院)

专题命中 其他图像生成 :image generation(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于正常逆伽玛分布的时间参数估计方法,用于长期阿尔茨海默病预测,通过生成中间图像和预测未来图像,保持疾病特征。

Comments 13pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏