arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-12-16 至 2025-12-16 共收录 7 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 7 篇

2510.09561 2025-12-16 cs.CV 79%

TC-LoRA: Temporally Modulated Conditional LoRA for Adaptive Diffusion Control

TC-LoRA: 基于时间调节的条件LoRA用于自适应扩散控制

Minkyoung Cho, Ruben Ohana, Christian Jacobsen, Adityan Jothi, Min-Hung Chen, Z. Morley Mao, Ethem Can

机构 * University of Michigan(密歇根大学) NVIDIA(英伟达)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 TC-LoRA通过动态调整模型权重实现自适应扩散控制,提升生成保真度和空间条件符合性。

Comments Project Page: https://minkyoungcho.github.io/tc-lora/; NeurIPS 2025 Workshop on SPACE in Vision, Language, and Embodied AI (SpaVLE); 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17364 2025-12-16 cs.CV cs.AI 79%

Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation

条件编织与专家调节:迈向通用且可控的图像生成

Guoqing Zhang, Xingtong Ge, Lu Shi, Xin Zhang, Muqing Xue, Wanru Xu, Yigang Cen, Yidong Li

机构 * State Key Laboratory of Advanced Rail Autonomous Operation(先进轨道交通自主运行国家重点实验室) School of Computer Science and Technology(计算机科学与技术学院) Visual Intellgence +X International Cooperation Joint Laboratory of MOE(教育部视觉智能+X国际合作联合实验室) Hong Kong University of Science and Technology(香港科技大学) SenseTime Research(商汤科技研究院) Beijing Jiaotong University(北京交通大学) SenseTime Research Institute(时光机器研究院)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

AI总结 提出UniGen框架,通过CoMoE模块和WeaveNet机制实现通用且可控的图像生成,提升效率和表达性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00227 2025-12-16 cs.CV cs.AI cs.RO 79%

Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes

Ctrl-Crash: 可控扩散用于逼真汽车碰撞

Anthony Gosselin, Ge Ya Luo, Luis Lara, Florian Golemo, Derek Nowrouzezahrai, Liam Paull, Alexia Jolicoeur-Martineau, Christopher Pal

机构 * McGill University(麦吉尔大学) CIFAR AI Chair(CIFAR人工智能 chair)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 Ctrl-Crash通过可控扩散生成逼真汽车碰撞视频,提升交通安全模拟的可控性和真实性。

Comments Under review at Pattern Recognition Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13014 2025-12-16 cs.CV 70%

JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation Promotion

JoDiffusion: 通过像素级标注联合扩散图像以促进语义分割

Haoyu Wang, Lei Zhang, Wenrui Liu, Dengyang Jiang, Wei Wei, Chen Ding

专题命中 可控生成 :image generation(abstract);diffusion(abstract);分类 cs.CV

AI总结 JoDiffusion通过联合扩散图像与像素级标注,提升语义分割的性能和可扩展性。

Comments Accepted at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13008 2025-12-16 cs.CV 57%

TWLR: Text-Guided Weakly-Supervised Lesion Localization and Severity Regression for Explainable Diabetic Retinopathy Grading

TWLR: 基于文本引导的弱监督病变定位与严重程度回归用于可解释性糖尿病视网膜病变分级

Xi Luo, Shixin Xu, Ying Xie, JianZhong Hu, Yuwei He, Yuhui Deng, Huaxiong Huang

机构 * Guangdong Provincial Key Laboratory of Interdisciplinary Research and Application for Data Science(广东省级交叉学科研究与数据科学应用重点实验室) Department of Statistics and Data Science, Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist大学统计与数据科学系) Faculty of Science, Hong Kong Baptist University(香港 Baptist大学科学学院) Data Science Research Center, Duke Kunshan University(杜克-昆山大学数据科学研究中心) Shanxi Provincial People’s Hospital(山西人民医院) The Fifth Clinical Medical school of Shanxi Medical University(山西医科大学第五临床医学院) Research Center for Mathematics, Beijing Normal University(北京师范大学数学研究中心) Department of Mathematics and Statistics, York University(约克大学数学与统计学系)

专题命中 可控生成 :inpainting(abstract);分类 cs.CV

AI总结 TWLR通过双阶段框架实现糖尿病视网膜病变的可解释性评估,结合视觉语言模型和弱监督分割,实现病变定位与严重程度回归。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12751 2025-12-16 cs.CV 57%

GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation

GenieDrive: 向具有物理意识的驾驶世界模型迈进:基于4D占用的视频生成

Zhenya Yang, Zhe Liu, Yuxiang Lu, Liping Hou, Chenxuan Miao, Siyi Peng, Bailan Feng, Xiang Bai, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Huazhong University of Science and Technology(华中科技大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

AI总结 GenieDrive通过4D占用引导的视频生成,实现物理意识的驾驶视频生成,提升预测精度和视频质量。

Comments The project page is available at https://huster-yzy.github.io/geniedrive_project_page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12664 2025-12-16 cs.CV 57%

InteracTalker: Prompt-Based Human-Object Interaction with Co-Speech Gesture Generation

InteracTalker: 基于提示的人-物体交互与同步手势生成

Sreehari Rajan, Kunal Bhosikar, Charu Sharma

机构 * Machine Learning Lab, IIIT Hyderabad, India(IIIT Hyderabad 机器学习实验室)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

AI总结 InteracTalker通过整合基于提示的物体感知交互与同步手势生成,实现了人-物体交互的统一框架,提升了动作的真实性和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏