arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 70159 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70159 篇

2603.26357 2026-04-07 cs.CV 79%

MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model

MPDiT:一种多块全局到局部变换器架构用于高效的流匹配和扩散模型

Quan Dao, Dimitris Metaxas

机构 * Rutgers University(罗格斯大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 MPDiT通过多块变换器架构在扩散和流匹配模型中实现高效计算,减少50%的GFLOPs消耗同时保持生成性能,改进时间与类别嵌入设计加速训练收敛。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.25073 2026-04-07 cs.CV 79%

GaMO: Geometry-aware Multi-view Diffusion Outpainting for Sparse-View 3D Reconstruction

GaMO:基于几何的多视角扩散补全用于稀疏视角3D重建

Yi-Chuan Huang, Hao-Jen Chien, Chin-Yang Lin, Ying-Huan Chen, Yu-Lun Liu

机构 * National Yang Ming Chiao Tung University(国立阳明交通大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出GaMO框架,通过多视角补全实现稀疏视角3D重建,无需训练即可在零样本条件下提升重建性能,实验表明其在稀疏视角下具有高效且准确的重建能力。

Comments Project page: https://yichuanh.github.io/GaMO/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23709 2026-04-07 cs.CV 79%

Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion

Stream-DiffVSR: 通过自回归扩散实现低延迟可流式传输视频超分辨率

Hau-Shiang Shiu, Chin-Yang Lin, Zhixiang Wang, Chi-Wei Hsiao, Po-Fan Yu, Yu-Chih Chen, Yu-Lun Liu

机构 * National Yang Ming Chiao Tung University(国立阳明交通大学) Shanda AI Research Tokyo(盛大AI研究东京) MediaTek Inc.(联发科技股份有限公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Stream-DiffVSR提出一种因果条件扩散框架,通过四步蒸馏去噪器、自回归时间引导模块和轻量时间感知解码器,实现高效在线视频超分辨率,显著降低延迟并提升质量。

Comments Project page: https://jamichss.github.io/stream-diffvsr-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17634 2026-04-07 cs.CV 79%

Efficient Score Pre-computation for Diffusion Models via Cross-Matrix Krylov Projection

通过交叉矩阵Krylov投影实现扩散模型的高效分数预计算

Kaikwan Lau, Andrew S. Na, Justin W. L. Wan

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种加速基于分数的扩散模型的新框架,通过将标准稳定扩散模型转换为Fokker-Planck形式,利用交叉矩阵Krylov投影方法在共享子空间中快速求解后续矩阵,实现15.8%至43.7%的计算时间减少,并在去噪任务中表现出更高的效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03206 2026-04-07 cs.LG cs.CV math.ST stat.ML stat.TH 79%

An Analytical Theory of Spectral Bias in the Learning Dynamics of Diffusion Models

扩散模型学习动力学中频谱偏置的分析理论

Binxu Wang, Cengiz Pehlevan

机构 * Kempner Institute, Harvard University(哈佛大学肯普纳研究所) SEAS, Harvard University(哈佛大学工程与应用科学学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出分析框架,揭示生成分布在扩散模型训练中的演变规律,发现频谱定律揭示高方差结构学习速度快于低方差细节,且局部卷积改变学习动态。

Comments 96 pages, 29 figures. Published in Advances in Neural Information Processing Systems, NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02847 2026-04-06 cs.CV 79%

HiDiGen: Hierarchical Diffusion for B-Rep Generation with Explicit Topological Constraints

HiDiGen:基于显式拓扑约束的层次扩散用于B-Rep生成

Shurui Liu, Weide Chen, Ancong Wu

机构 * School of Computer Science and Engineering, Sun Yat-sen University, China(中山大学计算机科学与工程学院) School of Intelligent Systems Engineering, Shenzhen Campus of Sun Yat-sen University, China(中山大学深圳校区智能工程学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 HiDiGen通过分阶段的拓扑约束生成B-Rep结构,结合层次扩散模型提升几何生成的多样性和拓扑有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02787 2026-04-06 cs.CV cs.AI 79%

LumaFlux: Lifting 8-Bit Worlds to HDR Reality with Physically-Guided Diffusion Transformers

LumaFlux:通过物理引导的扩散变压器将8位世界提升到HDR现实

Shreshth Saini, Hakan Gedik, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Google, Inc.(谷歌公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 LumaFlux通过物理和感知引导的扩散变压器实现8位SDR到10位HDR的重建,引入PGA、PCM和HDR Residual Coupler模块,提升亮度和色彩保真度,同时构建大规模训练数据集和评估基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02168 2026-04-03 cs.CV 79%

Reflection Generation for Composite Image Using Diffusion Model

基于扩散模型的复合图像反射生成

Haonan Zhao, Qingyang Liu, Jiaxuan Chen, Li Niu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出通过扩散模型生成复合图像中反射的方法,引入反射位置和外观先验信息,构建首个大规模物体反射数据集DEROBA,实现物理一致且视觉真实的反射生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01709 2026-04-03 cs.CV 79%

Bias mitigation in graph diffusion models

图扩散模型中的偏差缓解

Meng Yu, Kun Zhan

机构 * School of Information Science & Engineering, Lanzhou University(兰州大学信息科学与工程学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出方法缓解图扩散模型中的反向起始偏差和暴露偏差,通过改进的Levin采样算法和新的得分修正机制,无需修改网络即可在多个模型和任务中取得最佳效果。

Comments Accepted to ICLR 2025!

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22652 2026-04-03 cs.RO cs.CV 79%

Pixel Motion Diffusion is What We Need for Robot Control

我们为机器人控制需要像素运动扩散

E-Ro Nguyen, Yichi Zhang, Kanchana Ranasinghe, Xiang Li, Michael S. Ryoo

机构 * Stony Brook University(石溪大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DAWN框架通过结构化像素运动表示连接高层运动意图与底层机器人动作,实现端到端可训练系统,展示多任务性能和现实迁移能力。

Comments Accepted to CVPR 2026. Project page: https://eronguyen.github.io/DAWN

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01053 2026-04-02 cs.CV 79%

PHASOR: Anatomy- and Phase-Consistent Volumetric Diffusion for CT Virtual Contrast Enhancement

PHASOR:基于解剖和相位一致的体积扩散用于CT虚拟对比增强

Zilong Li, Dongyang Li, Chenglong Ma, Zhan Feng, Dakai Jin, Junping Zhang, Hao Luo, Fan Wang, Hongming Shan

机构 * Shanghai Key Lab of Intelligent Information Processing, School of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院上海市智能信息处理重点实验室) Institute of Science and Technology for Brain-inspired Intelligence and MOE Frontiers Center for Brain Science, Fudan University(复旦大学类脑智能科学与技术研究院及教育部脑科学前沿中心) Shanghai Center for Brain Science and Brain-inspired Technology(上海脑科学与类脑研究中心) Department of Radiology, First Affiliated Hospital of College of Medical Science, Zhejiang University(浙江大学医学院附属第一医院放射科) Alibaba DAMO Academy(阿里巴巴达摩院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出PHASOR框架,通过体积扩散模型提升CT虚拟对比增强质量,引入解剖路由混合专家和强度-相位意识表示对齐模块,解决解剖异质性和空间偏移问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00985 2026-04-02 cs.CV 79%

Maximizing T2-Only Prostate Cancer Localization from Expected Diffusion Weighted Imaging

从预期扩散加权成像最大化T2-only前列腺癌定位

Weixi Yi, Yipei Wang, Wen Yan, Hanyuan Zhang, Natasha Thorley, Alexander Ng, Shonit Punwani, Fernando Bianco, Mark Emberton, Veeru Kasivisvanathan, Dean C. Barratt, Shaheer U. Saeed, Yipeng Hu

机构 * UCL Hawkes Institute and the Department of Medical Physics and Biomedical Engineering, University College London(伦敦大学学院霍克斯研究所与医学物理与生物医学工程系) Centre for Bioengineering, School of Engineering and Materials Science and Digital Environment Research Institute, Queen Mary University of London(伦敦玛丽女王大学工程与材料科学学院生物工程中心与数字环境研究所) Centre for Medical Imaging, University College London(伦敦大学学院医学影像中心) Division of Surgery & Interventional Science, University College London(伦敦大学学院外科与介入科学部) Urological Research Network(泌尿研究网络)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出利用T2加权图像和生成模型进行前列腺癌定位,通过期望最大化算法提升定位性能,相比传统方法在患者和区域层面表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00853 2026-04-02 cs.CV 79%

MotionGrounder: Grounded Multi-Object Motion Transfer via Diffusion Transformer

MotionGrounder: 通过扩散变换器实现多物体运动迁移

Samuel Teodoro, Yun Chen, Agus Gunawan, Soo Ye Kim, Jihyong Oh, Munchurl Kim

机构 * School of Electrical Engineering, Korea Advanced Institute of Science and Technology(韩国科学技术院电气工程学院) Adobe Research(Adobe研究院) CMLab, Chung-Ang University(中央大学CMLab)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出MotionGrounder,一种基于扩散变换器的多物体运动迁移框架,通过流式运动信号和对象-描述对齐损失实现稳定运动先验和空间对齐,实验显示其在定量、定性和人类评估中均优于现有方法。

Comments Please visit our project page at https://kaist-viclab.github.io/motiongrounder-site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00519 2026-04-02 cs.CV 79%

Learnability-Guided Diffusion for Dataset Distillation

基于可学习性的扩散模型用于数据集蒸馏

Jeffrey A. Chan-Santiago, Mubarak Shah

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出基于可学习性的数据集蒸馏方法,通过逐步构建合成数据集,减少冗余信号,提升模型性能,在ImageNet-1K等数据集上取得最佳结果。

Comments This paper has been accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29742 2026-04-02 cs.CV cs.CR 79%

SHIFT: Stochastic Hidden-Trajectory Deflection for Removing Diffusion-based Watermark

SHIFT: 随机隐藏轨迹偏移用于移除基于扩散的水印

Rui Bao, Zheng Gao, Xiaoyu Li, Xiaoyan Feng, Yang Song, Jiaojiao Jiang

机构 * University of New South Wales(新南威尔士大学) Griffith University(格里菲斯大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 SHIFT通过随机隐藏轨迹偏移攻击利用扩散水印的轨迹恢复漏洞,实现高成功率的水印移除,不依赖模型训练,保持视觉和语义一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14464 2026-04-02 cs.CV cs.AI 79%

CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style Transfer

CoCoDiff:基于语义对应关系的扩散模型用于细粒度风格迁移

Wenbo Nie, Zixiang Li, Renshuai Tao, Bin Wu, Yunchao Wei, Yao Zhao

机构 * Institute of Information Science, Beijing Jiaotong University(北京交通大学信息科学研究所) Visual Intelligence + X International Joint Laboratory of the Ministry of Education(视觉智能+X教育部国际联合实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 CoCoDiff提出一种无需训练的低成本风格迁移框架,利用预训练的潜在扩散模型实现细粒度语义一致的风格化,通过像素级语义对应模块和循环一致性模块实现对象和区域级别的风格迁移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16148 2026-04-02 cs.CV 79%

ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion

ActionMesh: 带有时间轴的3D扩散模型用于动画3D网格生成

Remy Sabathier, David Novotny, Niloy J. Mitra, Tom Monnier

机构 * Meta Reality Labs(Meta现实实验室) SpAItial University College London(伦敦大学学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 ActionMesh通过引入时间轴的3D扩散模型,实现了从多种输入生成动画3D网格,具备快速生成、无骨骼和拓扑一致性的优势,提升了在纹理和重定向等应用中的效率。

Comments CVPR 2026. Project webpage with code and videos: https://remysabathier.github.io/actionmesh/ . V2 update includes more baseline models with a larger evaluation set on our new publicly released benchmark ActionBench, and {3D+video}-to-animated-mesh qualitative comparison in supplemental

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00250 2026-04-02 cs.CV 79%

PRISM: Differentiable Analysis-by-Synthesis for Fixel Recovery in Diffusion MRI

PRISM:基于微结构拟合的固定点恢复分析-合成可微方法

Mohamed Abouagour, Atharva Shah, Eleftherios Garyfallidis

机构 * Indiana University, Bloomington, IN, USA(印第安纳大学布卢明顿分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 PRISM通过端到端的空间块拟合显式多室模型,结合CSF、灰质和白质纤维方向,利用排斥和稀疏先验实现纤维恢复,优于现有方法。

Comments 10 pages, 1 figure, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00093 2026-04-02 cs.CV 79%

RawGen: Learning Camera Raw Image Generation

RawGen: 学习相机原始图像生成

Dongyoung Kim, Junyong Lee, Abhijith Punnappurath, Mahmoud Afifi, Sangmin Han, Alex Levinshtein, Michael S. Brown

机构 * Samsung Electronics(三星电子) Yonsei University(延世大学)

专题命中 扩散模型 :image generation(title);diffusion(abstract);分类 cs.CV

AI总结 RawGen通过扩散模型生成任意目标相机的原始图像,解决大规模原始数据集收集困难的问题,并提升低级视觉任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29569 2026-04-01 cs.GR 79%

AdaptDiff: Adaptive Guidance in Diffusion Models for Diverse and Identity-Consistent Face Synthesis (Student Abstract)

AdaptDiff: 差分模型中适应性引导以生成多样且身份一致的面部合成

Eduarda Caldeira, Tahar Chettaoui, Naser Damer, Fadi Boutros

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.GR

AI总结 本文提出动态加权方案,通过适应性引导抑制无关属性,提升生成面部的一致性和多样性。

Comments Accepted at AAAI 2026 Student Abstract and Poster Program

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15713 2026-04-01 cs.CV 79%

DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models

DiffusionVL: 将任意自回归模型转化为扩散视觉语言模型

Lunbin Zeng, Jingfeng Yao, Bencheng Liao, Hongyuan Tao, Wenyu Liu, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出DiffusionVL,通过高效扩散微调将预训练自回归模型转化为扩散视觉语言模型,实现高效生成和性能提升。

Comments 12 pages, 4 figures, conference or other essential info

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19114 2026-04-01 cs.CV 79%

CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design

CreatiDesign: 一种统一的多条件扩散变换器用于创意图形设计

Hui Zhang, Dexiang Hong, Maoke Yang, Yutao Cheng, Zhao Zhang, Weidong Chen, Jie Shao, Xinglong Wu, Zuxuan Wu, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究所) Shanghai Key Laboratory of Multimodal Embodied AI(上海市多模态具身人工智能重点实验室) Bytedance Intelligent Creation(字节跳动智能创作) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学技术学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出CreatiDesign,一种统一的多条件扩散变换器,解决多条件控制问题,通过多模态注意力掩码机制和自动化数据集构建,提升图形设计生成的准确性和和谐度。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16205 2026-04-01 cs.CV eess.IV 79%

Generative AI Enables Structural Brain Network Construction from fMRI via Symmetric Diffusion Learning

生成式AI通过对称扩散学习从fMRI构建结构脑网络

Qiankun Zuo, Bangjun Lei, Wanyu Qiu, Changhong Jing, Jin Hong, Shuqiang Wang

机构 * Department of Computing, Hong Kong Polytechnic University(香港理工大学计算学系) School of Information Engineering, Nanchang University(南昌大学信息工程学院) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出DiffGAN-F2S模型,通过统一框架从fMRI预测结构连接,利用对称扩散生成对抗网络和自适应学习生成高保真结构连接,测试表明其在结构连接预测上优于其他模型,并能识别重要脑区和连接,为多模态脑网络融合提供新方法。

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29343 2026-04-01 cs.CV 79%

FOSCU: Feasibility of Synthetic MRI Generation via Duo-Diffusion Models for Enhancement of 3D U-Nets in Hepatic Segmentation

FOSCU:通过双扩散模型生成合成MRI的可行性,以增强肝部分割的3D U-Net

Youngung Han, Kyeonghun Kim, Seoyoung Ju, Yeonju Jean, Minkyung Cha, Seohyoung Park, Hyeonseok Jung, Nam-Joon Kim, Woo Kyoung Jeong, Ken Ying-Kai Liao, Hyuk-Jae Lee

机构 * Seoul National University(首尔国立大学) Sangmyung University(祥明大学) Ewha Womans University(梨花女子大学) Chung-Ang University(中央大学) Samsung Medical Center, Sungkyunkwan University School of Medicine(三星医学中心,成均馆大学医学院) NVIDIA(英伟达)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出FOSCU,通过双扩散模型生成高分辨率合成MRI数据及分割标签,结合改进的3D U-Net训练流程,提升肝部分割的鲁棒性与图像真实性。

Comments 10 pages, 5 figures. Accepted at IEEE APCCAS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15968 2026-04-01 cs.CV 79%

HyperAlign: Hypernetwork for Efficient Test-Time Alignment of Diffusion Models

HyperAlign:用于扩散模型高效测试时对齐的超网络

Xin Xie, Jiaxian Guo, Dong Gong

机构 * University of New South Wales (UNSW Sydney)(新南威尔士大学(悉尼新南威尔士大学)) Google Research(谷歌研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 HyperAlign通过训练超网络实现扩散模型测试时对齐,平衡对齐质量与计算效率,提升语义一致性和视觉质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24176 2026-04-01 cs.CV cs.LG 79%

Guiding a Diffusion Transformer with the Internal Dynamics of Itself

通过自身内部动态引导扩散变换器

Xingyu Zhou, Qifan Li, Xiaobin Hu, Hai Chen, Shuhang Gu

机构 * University of Electronic Science and Technology of China(电子科技大学) National University of Singapore(新加坡国立大学) Sun Yat-sen University(中山大学) North China Institute of Computer Systems Engineering(华北计算机系统工程研究所)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出内部引导策略,通过训练过程中中间层的辅助监督和采样时的中间层输出扩展,提升训练效率和生成质量,在ImageNet上取得显著效果。

Comments Accepted to CVPR 2026. Project Page: https://zhouxingyu13.github.io/Internal-Guidance/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11423 2026-04-01 cs.CV 79%

JoyStreamer-Flash: Real-time and Infinite Audio-Driven Avatar Generation with Autoregressive Diffusion

JoyStreamer-Flash: 基于自回归扩散的实时和无限音频驱动化身生成

Chaochao Li, Ruikui Wang, Liangbo Zhou, Jinheng Feng, Huaishao Luo, Huan Zhang, Youzheng Wu, Xiaodong He

机构 * JD Technology(京东科技)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出JoyStreamer-Flash,一种能实现实时推断和无限长度视频生成的音频驱动自回归模型,通过渐进式步进 bootstrap、运动条件注入和无界RoPE via Cache-Resetting提升生成质量与一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04466 2026-04-01 cs.CV 79%

Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion

基于解耦潜在扩散的情感语音驱动3D人体动画

Kiran Chhatre, Radek Daněček, Nikos Athanasiou, Giorgio Becherini, Christopher Peters, Michael J. Black, Timo Bolkart

机构 * KTH Royal Institute of Technology, Sweden(瑞典皇家理工学院) Max Planck Institute for Intelligent Systems, Germany(德国马克斯·普朗克智能系统研究所)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出AMUSE模型,通过解耦潜在向量实现情感语音驱动的3D人体动画生成,能控制情绪和风格,生成更符合语音内容且表现更真实的情绪动作序列。

Comments Conference on Computer Vision and Pattern Recognition (CVPR) 2024. Webpage: https://amuse.is.tue.mpg.de/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28114 2026-03-31 cs.CV cs.LG 79%

Attention Frequency Modulation: Training-Free Spectral Modulation of Diffusion Cross-Attention

注意力频率调制:无训练的扩散交叉注意力频谱调制

Seunghun Oh, Unsang Park

机构 * Sogang University(西江大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出AFM,通过频域调整交叉注意力logits,实现无训练的注意力频谱控制,实验证明能有效编辑图像内容并保持语义一致性。

Comments 16 pages; preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26068 2026-03-31 cs.CV 79%

PAD-Hand: Physics-Aware Diffusion for Hand Motion Recovery

PAD-Hand: 基于物理的扩散方法用于手部运动恢复

Elkhan Ismayilzada, Yufei Zhang, Zijun Cui

机构 * Michigan State University(密歇根州立大学) Independent Researcher(独立研究者)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种物理感知的条件扩散框架,通过MeshCNN-Transformer架构,利用欧拉-拉格朗日动力学提升手部运动的物理一致性,并通过方差估计增强运动的可解释性。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏