arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 70159 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70159 篇

2604.23651 2026-04-28 cs.CV 79%

Geometry-Conditioned Diffusion for Occlusion-Robust In-Bed Pose Estimation

基于几何条件的扩散模型用于抗遮挡的卧床姿态估计

Navid Aslankhani Khameneh, Marco Carletti, Cigdem Beyan

机构 * Department of Computer Science, University of Verona(威尼斯大学计算机科学系) EVS - Embedded Vision Systems Srl(嵌入式视觉系统股份有限公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出基于几何条件的扩散模型,解决卧床姿态估计中遮挡问题,通过生成模型直接从骨骼关键点生成遮挡图像,提升遮挡鲁棒性。

Comments This is the preprint version of the paper. The final version has been accepted for publication in the Proceedings of the 20th IEEE International Conference on Automatic Face and Gesture Recognition (IEEE FG 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23636 2026-04-28 cs.CV 79%

Discriminator-Guided Adaptive Diffusion for Source-Free Test-Time Adaptation under Image Corruptions

判别器引导的自适应扩散用于无源测试时间适应下的图像损坏

Francesco Olivato, Cigdem Beyan, Vittorio Murino

机构 * Department of Computer Science, University of Verona(威尼斯大学计算机科学系) AI for Good (AIGO), Istituto Italiano di Tecnologia(意大利技术研究院人工智能与善部(AIGO))

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种基于扩散的输入级适应框架,在测试时保持所有源训练模型冻结,通过判别器引导的自适应扩散策略动态控制每个测试样本的扰动量,以抑制特定域的损坏,提升鲁棒性。

Comments This is the preprint (submitted version) of the paper. The final version has been accepted for publication in the Proceedings of the 28th International Conference on Pattern Recognition (ICPR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04749 2026-04-28 cs.CV 79%

Mitigating Long-Tail Bias via Prompt-Controlled Diffusion Augmentation

通过提示控制的扩散增强缓解长尾偏差

Buddhi Wijenayake, Nichula Wasalathilake, Roshan Godaliyadda, Vijitha Herath, Parakrama Ekanayake, Vishal M. Patel

机构 * University of Peradeniya(珀德尼亚大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出基于提示控制的扩散增强框架,通过生成带标签的图像样本来增强少数类,提升遥感图像分割中长尾不平衡问题的处理效果。

Comments Accepted to Publication at 2026 IEEE International Geoscience and Remote Sensing Symposium (IGARSS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06904 2026-04-28 cs.CV 79%

BIR-Adapter: A parameter-efficient diffusion adapter for blind image restoration

BIR-Adapter:一种参数高效的扩散适配器用于盲图像恢复

Cem Eteke, Alexander Griessel, Wolfgang Kellerer, Eckehard Steinbach

机构 * Chair of Media Technology, Munich Institute of Robotics and Machine Intelligence(媒体技术教授会,慕尼黑机器人与机器智能研究所) School of Computation, Information and Technology, Technical University of Munich(计算、信息与技术学院,慕尼黑技术大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出BIR-Adapter,一种参数高效的扩散适配器,用于盲图像恢复。通过减少训练参数数量和引入采样引导机制,提升恢复可靠性,实验表明其在多个设置中性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22808 2026-04-28 cs.CV cs.AI eess.IV 79%

FreqFormer: Hierarchical Frequency-Domain Attention with Adaptive Spectral Routing for Long-Sequence Video Diffusion Transformers

FreqFormer:具有自适应频谱路由的分层频域注意力机制用于长序列视频扩散变换器

Haopeng Jin

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FreqFormer通过分层频域注意力机制,利用视频特征的频谱结构,采用不同操作符处理不同频段,降低长序列视频扩散变换器的计算与内存开销。

Comments 24 pages, 17 figures, 14 tables, Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21066 2026-04-28 cs.CV cs.LG stat.ME 79%

Optimizing Diffusion Priors in Image Reconstruction from a Single Observation

从单个观测重建图像中优化扩散先验

Frederic Wang, Katherine L. Bouman

机构 * Caltech(加州理工学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出通过结合现有扩散先验生成单专家先验并优化指数,以提升单观测图像重建的可靠性,验证了在黑洞成像和文本条件去模糊中效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18836 2026-04-28 eess.IV cs.AI cs.CV 79%

Dual-domain Multi-path Self-supervised Diffusion Model for Accelerated MRI Reconstruction

双域多路径自监督扩散模型用于加速MRI重建

Yuxuan Zhang, Jinkui Hao, Bo Zhou

机构 * Department of Radiology, Northwestern University(放射科,西北大学) Department of Biomedical Engineering, Huazhong University of Science and Technology(生物医学工程系,华中科技大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出双域多路径自监督扩散模型,通过自监督训练方案、轻量混合注意力网络和多路径推理策略,提升MRI重建的准确性、效率和可解释性,克服传统模型依赖全采样数据的局限。

Comments Accepted at IEEE-TNNLS, 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22220 2026-04-27 cs.CV 79%

Breaking Watermarks in the Frequency Domain: A Modulated Diffusion Attack Framework

频域中破解水印:一种调制扩散攻击框架

Chunpeng Wang, Binyan Qu, Xiaoyu Wang, Zhiqiu Xia, Shanshan Zhang, Yunan Liu, Qi Li

机构 * Qilu University of Technology (Shandong Academy of Sciences)(齐鲁工业大学(山东省科学院)) Dalian Maritime University(大连海事大学) Nanjing University of Science and Technology(南京理工大学) Shandong Jianzhu University(山东建筑大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出FMDiffWA框架,通过频域调制模块在正反扩散过程中实现水印信号的高效去除,同时保持图像视觉质量,改进了传统扩散模型的训练策略,实验显示其在视觉保真度和通用性方面优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22212 2026-04-27 eess.IV cs.CV cs.LG 79%

Multimodal Diffusion to Mutually Enhance Polarized Light and Low Resolution EBSD Data

多模态扩散用于互为增强的偏振光与低分辨率EBSD数据

Harry Dong, Timofey Efimov, Megna Shah, Jeff Simmons, Sean Donegan, Marc De Graef, Yuejie Chi

机构 * Carnegie Mellon University(卡内基梅隆大学) Air Force Research Laboratory(空军研究实验室) Yale University(耶鲁大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出多模态扩散模型,通过结合偏振光和低分辨率EBSD数据,提升数据收集效率,实现颗粒边界预测、超分辨率和去噪等任务的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20886 2026-04-27 cs.CV cs.LG eess.IV 79%

Nuclear Diffusion Models for Low-Rank Background Suppression in Videos

核扩散模型用于视频中低秩背景抑制

Tristan S. W. Stevens, Oisín Nolan, Jean-Luc Robert, Ruud J. G. van Sloun

机构 * Dept. of Electrical Engineering, Eindhoven University of Technology, the Netherlands(埃因霍温理工大学电气工程系,荷兰) Philips Research North America, Cambridge MA, USA(飞利浦北美研究部,马萨诸塞州剑桥)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出核扩散模型,结合低秩时间建模与扩散后验采样,提升视频去雾性能,尤其在对比度增强和信号保持方面优于传统RPCA。

Comments 5 pages, 4 figures, preprint

Journal ref 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21592 2026-04-24 cs.CV 79%

Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers

Sculpt4D:通过稀疏注意力扩散变换器生成4D形状

Minghao Yin, Wenbo Hu, Jiale Xu, Ying Shan, Kai Han

机构 * The University of Hong Kong(香港大学) ARC Lab, Tencent PCG(腾讯PCG实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Sculpt4D通过在预训练的3D扩散变换器中集成高效时间建模,解决4D生成中的时间伪影和计算需求问题,实现高保真4D合成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21518 2026-04-24 eess.IV cs.CV 79%

DiffNR: Diffusion-Enhanced Neural Representation Optimization for Sparse-View 3D Tomographic Reconstruction

DiffNR: 用于稀疏视图3D断层成像重建的扩散增强神经表示优化

Shiyan Su, Ruyi Zha, Danli Shi, Hongdong Li, Xuelian Cheng

机构 * Monash University(莫纳什大学) The Australian National University(澳大利亚国立大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DiffNR通过引入扩散先验改进神经表示优化,解决稀疏视图下CT成像的严重伪影问题,提升PSNR并保持高效优化。

Comments Accepted to AAAI 2026. Project page: https://ooonesevennn.github.io/DiffNR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21422 2026-04-24 cs.CV 79%

Pre-process for segmentation task with nonlinear diffusion filters

分割任务前的非线性扩散滤波预处理

Javier Sanguino, Carlos Platero, Olga Velasco

机构 * Health Science Technology Group, Technical University of Madrid(健康科学技术组,马德里技术大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出非线性扩散滤波用于分割前的预处理,设计新型扩散函数以实现清晰轮廓和无模糊边缘,证明滤波器满足尺度空间要求,通过半隐式方案高效生成分段常数图像。

Comments Manuscript from 2017, previously unpublished, 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21279 2026-04-24 cs.CV 79%

LatRef-Diff: Latent and Reference-Guided Diffusion for Facial Attribute Editing and Style Manipulation

LatRef-Diff: 基于潜在和参考引导的扩散模型用于面部属性编辑和风格操控

Wenmin Huang, Weiqi Luo, Xiaochun Cao, Jiwu Huang

机构 * GuangDong Province Key Lab of Information Security Technology and School of Computer Science and Engineering(广东信息安全技术重点实验室和计算机科学与工程学院) School of Cyber Science and Technology, Shenzhen Campus(深圳校区网络科学与技术学院) Guangdong Laboratory of Machine Perception and Intelligent Computing, Faculty of Engineering(广东机器感知与智能计算实验室,工程学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出LatRef-Diff,通过潜在和参考引导方法改进面部属性编辑和风格操控,解决传统扩散模型在属性控制和风格操控中的不足,采用风格代码和模块化设计提升精度和图像质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21221 2026-04-24 cs.CV cs.LG 79%

Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation

稀疏强制:用于实时自回归扩散视频生成的原生可训练稀疏注意力

Boxun Xu, Yuming Du, Zichang Liu, Siyu Yang, Ziyang Jiang, Siqi Yan, Rajasi Saha, Albert Pumarola, Wenchen Wang, Peng Li

机构 * Meta Superintelligence Labs(Meta超智能实验室) University of California, Santa Barbara(加州大学圣芭芭拉分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出稀疏强制方法,通过在训练和推理中提升长时程生成质量并降低解码延迟,采用可训练的原生稀疏机制和高效的GPU内核PBSA来加速稀疏注意力和内存更新。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17656 2026-04-24 cs.SD cs.AI cs.CL cs.CV cs.LG 79%

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

Video-Robin:基于意图的视频到音乐生成的自回归扩散规划

Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj, Gouthaman KV, Ramani Duraiswami, Lie Lu, Sreyan Ghosh, Dinesh Manocha

机构 * University of Maryland College Park(马里兰大学College Park分校) Dolby Laboratories(杜比实验室) NVIDIA(英伟达)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Video-Robin通过结合自回归规划与扩散合成,实现高质量的语义对齐音乐生成,相比传统方法在速度和质量上均有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18457 2026-04-24 cs.CV cs.LG 79%

VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models

VFM-VAE:视觉基础模型可以作为潜在扩散模型的良好分词器

Tianci Bi, Xiaoyi Zhang, Yan Lu, Nanning Zheng

机构 * State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人机混合增强智能国家重点实验室,人工智能与机器人研究院,西安交通大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出VFM-VAE,利用冻结的视觉基础模型作为潜在扩散模型的分词器,通过设计新解码器提升图像重建能力,实现分词器与扩散模型的协同优化,提升训练效率与性能。

Comments Accepted at CVPR 2026. Code and models available at: https://github.com/tianciB/VFM-VAE

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13576 2026-04-24 eess.IV cs.CV 79%

Cross-Distribution Diffusion Priors-Driven Iterative Reconstruction for Sparse-View CT

跨分布扩散先验驱动的稀疏视图CT迭代重建

Haodong Li, Shuo Han, Haiyang Mao, Yu Shi, Changsheng Fang, Jianjia Zhang, Weiwen Wu, Hengyong Yu

机构 * Department of Electrical and Computer Engineering, University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校电子与计算机工程系) School of Biomedical Engineering, Sun Yat-Sen University(中山大学生物医学工程学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出CDPIR框架,通过整合跨分布扩散先验与模型驱动迭代重建方法,解决稀疏视图CT在分布外场景下的重建问题,提升图像质量与泛化能力。

Comments 17 pages, 15 figures, accepted by IEEE Transactions on Medical Imaging

Journal ref IEEE Transactions on Medical Imaging, 2026 (early access)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20470 2026-04-23 cs.CV 79%

DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion

动态Rad:面向长视频扩散的内容自适应稀疏注意力

Yongji Long, Shijun Liang, Jintao Li, Yun Li

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Shenzhen Institute for Advanced Study, UESTC(深圳先进研究院) Computational Mathematics, Science, & Engineering at Michigan State University (MSU)(密歇根州立大学计算数学、科学与工程)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出DynamicRad,通过引入双模式策略和语义运动路由器,实现长视频扩散中的高效稀疏注意力机制,提升推理速度并保持高质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20715 2026-04-23 cs.CV 79%

GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers

GeoRelight: 基于灵活多模态扩散变换器的联合几何重照明与重建

Yuxuan Xue, Ruofan Liang, Egor Zakharov, Timur Bagautdinov, Chen Cao, Giljoo Nam, Shunsuke Saito, Gerard Pons-Moll, Javier Romero

机构 * Codec Avatars Lab, Meta(Meta编码器动画实验室) University of Tübingen(图宾根大学) Max Planck Institute for Informatics, Saarland Informatics Campus(马克斯·普朗克信息学院,萨尔兰信息校园)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出GeoRelight,通过联合解决几何重建与重照明问题,利用iNOD和混合数据训练方法提升性能,优于传统顺序模型和忽略几何的系统。

Comments CVPR 2026 Highlight; Project page: https://yuxuan-xue.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20594 2026-04-23 cs.CV 79%

Physics-Informed Conditional Diffusion for Motion-Robust Retinal Temporal Laser Speckle Contrast Imaging

具有物理信息的条件扩散用于运动鲁棒的视网膜时间激光散斑对比成像

Qian Chen, Yuehao Chen, Qiang Wang, Lei Zhu, Yanye Lu, Qiushi Ren

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出RetinaDiff框架,通过相位相关注册和条件扩散模型实现运动鲁棒的视网膜时间激光散斑对比成像,提升有限帧下的结构连续性和统计稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19747 2026-04-22 cs.CV 79%

AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model

AnyRecon:基于视频扩散模型的任意视角3D重建

Yutian Chen, Shi Guo, Renbiao Jin, Tianshuo Yang, Xin Cai, Yawen Luo, Mingxin Yang, Mulin Yu, Linning Xu, Tianfan Xue

机构 * AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model(AnyRecon:任意视角3D重建与视频扩散模型)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 AnyRecon通过视频扩散模型实现任意视角3D重建,保留显式几何控制并支持灵活条件设置,通过全局场景记忆和几何感知条件策略提升大场景重建效率与鲁棒性。

Comments Webpage: https://yutian10.github.io/AnyRecon/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19392 2026-04-22 cs.CV 79%

HarmoniDiff-RS: Training-Free Diffusion Harmonization for Satellite Image Composition

HarmoniDiff-RS:无训练扩散谐调用于卫星图像合成

Xiaoqi Zhuang, Jefersson A. Dos Santos, Jungong Han

机构 * The University of Sheffield(谢菲尔德大学) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出HarmoniDiff-RS,一种无训练的扩散框架,用于在不同领域条件下谐调合成卫星图像。通过潜在均值偏移操作对齐源域和目标域,结合时间步的潜在融合策略生成候选合成图像,并利用轻量级和谐分类器选择最一致的结果。

Comments 8 pages, 6 figures, CVPR 2026 findings. Code is available at https://github.com/XiaoqiZhuang/HarmoniDiff-RS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04818 2026-04-22 cs.CV eess.IV stat.ML 79%

Single-Step Reconstruction-Free Anomaly Detection and Segmentation via Diffusion Models

基于扩散模型的实时无重建异常检测与分割

Mehrdad Moradi, Marco Grasso, Bianca Maria Colosimo, Kamran Paynabar

机构 * H. Milton Stewart School of Industrial and Systems Engineering(H. Milton Stewart工业与系统工程学院) Georgia Institute of Technology(佐治亚理工学院) Department of Mechanical Engineering(机械工程系) Polytechnic University of Milan(米兰理工学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出RADAR方法,通过注意力机制的扩散模型直接生成异常图,提升检测精度和效率,实验证明在MVTec-AD和3D打印材料数据集上均优于现有方法。

Comments 9 pages, 8 figures, 1 table. Accepted to 2025 International Conference on Machine Learning and Applications (ICMLA)

Journal ref Proc. 2025 International Conference on Machine Learning and Applications (ICMLA), Boca Raton, FL, USA, 2025, pp. 663-670

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19105 2026-04-22 cs.CV 79%

EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation

视角运动:层次推理与扩散用于中心视觉-语言运动生成

Ruibing Hou, Mingyue Zhou, Yuwei Gui, Mingshuang Luo, Bingpeng Ma, Hong Chang, Shiguang Shan, Xilin Chen

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS)(智能信息处理重点实验室,计算技术研究所(ICT),中国科学院(CAS)) Jilin University (JLU)(吉林大学(JLU)) Beijing University of Posts and Telecommunications (BUPT)(北京邮电大学(BUPT)) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出EgoMotion框架,通过层次推理和扩散模型生成基于第一视角视觉和语言指令的3D人类运动,解决语义推理与运动建模的协同优化问题,提升多模态接地和运动质量。

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18632 2026-04-22 cs.CV stat.AP 79%

StomaD2: An All-in-One System for Intelligent Stomatal Phenotype Analysis via Diffusion-Based Restoration Detection Network

StomaD2:一种基于扩散恢复检测网络的智能气孔表型分析一体化系统

Quanling Zhao, Meng'en Qin, Yanfeng Sun, Yuan Miao, Xiaohui Yang

机构 * Henan Engineering Research Center for Artificial Intelligence Theory and Algorithms(河南人工智能理论与算法工程研究中心) School of Mathematics and Statistics(数学与统计学学院) International Joint Research Laboratory for Global Change Ecology(全球变化生态学联合研究实验室) School of Life Sciences(生命科学学院) State Key Laboratory of Cotton Biology(棉花生物学国家重点实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 StomaD2通过扩散恢复模块和专用旋转目标检测网络,实现复杂成像条件下高精度快速气孔表型分析,其柱状结构、上下文感知重采样机制和特征重组模块提升了特征表示能力,实验表明其在玉米和小麦数据集上准确率达99.4%和99.2%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09806 2026-04-22 cs.CV cs.AI cs.CL cs.LG 79%

A Generalist Model for Diverse Text-Guided Medical Image Synthesis

一种通用模型用于多样化文本引导的医学图像合成

Joseph Cho, Mrudang Mathur, Cyril Zakka, Dhamanpreet Kaur, Matthew Leipzig, Alex Dalal, Aravind Krishnan, Eubee Koo, Karen Wai, Cindy S. Zhao, Akshay Chaudhari, Matthew Duda, Ashley Choi, Ehsan Rahimy, Lyna Azzouz, Robyn Fong, Rohan Shad, William Hiesinger

机构 * Department of Cardiothoracic Surgery(心脏外科部门) Stanford Medicine(斯坦福医学) Department of Ophthalmology(眼科部门) Division of Cardiovascular Surgery(心血管外科分会) Penn Medicine(宾夕法尼亚医学)

专题命中 扩散模型 :image synthesis(title);diffusion(abstract);分类 cs.CV

AI总结 本文提出MediSyn模型,通过公开数据训练,生成六种医学领域和十种成像模态的合成图像,验证了通用模型在效率、质量和医学应用中的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18393 2026-04-21 cs.CV 79%

One-Step Diffusion with Inverse Residual Fields for Unsupervised Industrial Anomaly Detection

一步扩散与逆残差场用于无监督工业异常检测

Boan Zhang, Wen Li, Guanhua Yu, Xiyang Liu, Wenchao Chen, Long Tian

机构 * Department of Computer Science and Technology(计算机科学与技术系) Xidian University(西电大学) Department of Electronic Engineering(电子工程系)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出OSD-IRF方法,通过逆残差场实现无监督工业异常检测,利用单步扩散提升推理效率,并在多个基准测试中取得最优或竞争性性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18313 2026-04-21 cs.CV 79%

Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection

去噪与对齐:基于扩散的前景知识提示用于开放词汇时序动作检测

Sa Zhu, Wanqian Zhang, Lin Wang, Jinchao Zhang, Cong Wang, Bo Li

机构 * Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences State Key Laboratory of Cyberspace Security Defense Beijing China Institute of Information Engineering, Chinese Academy of Sciences Beijing China Hangzhou Dianzi University Hangzhou China Institute of Information Engineering, Chinese Academy of Sciences\ Key Laboratory of Cyberspace Security Defense Beijing China Engineering, Zhejiang University Hangzhou China Institute of Information Engineering, Chinese Academy of Sciences State Key Laboratory of Cyberspace Security Defense Beijing China Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences State Key Laboratory of Cyberspace Security Defense Institute of Information Engineering, Chinese Academy of Sciences Hangzhou Dianzi University Institute of Information Engineering, Chinese Academy of Sciences\ Key Laboratory of Cyberspace Security Defense Engineering, Zhejiang University Institute of Information Engineering, Chinese Academy of Sciences State Key Laboratory of Cyberspace Security Defense

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出DFAlign框架,通过扩散去噪生成前景知识,解决开放词汇时序动作检测中语义不平衡问题,提升动作相关片段的判别性。

Comments Accepted by SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13182 2026-04-21 cs.CV 79%

Diffusion-Based Feature Denoising and Using NNMF for Robust Brain Tumor Classification

基于扩散的特征去噪和使用NNMF进行鲁棒性脑肿瘤分类

Hiba Adil Al-kharsan, Róbert Rajkó

机构 * Doctoral School of Computer Science, University of Szeged(计算机科学博士学院,塞格德大学) Academic Staff, Doctoral School of Computer Science, University of Szeged(学术人员,计算机科学博士学院,塞格德大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出结合NNMF、轻量CNN和扩散去噪的鲁棒脑肿瘤分类框架,通过预处理MRI图像并提取可解释特征,提升对抗扰动下的分类鲁棒性。

Comments 30 pages, 29 figures

Journal ref Mach. Learn. Knowl. Extr. 2026, 8(4), 105

详情

展开后加载摘要…

URL PDF HTML 收藏