arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86714 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70159 篇

2503.18719 2026-01-08 cs.CV 83%

Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings

通过随机化位置编码提升扩散变换器的分辨率泛化能力

Liang Hou, Cong Liu, Mingwu Zheng, Xin Tao, Pengfei Wan, Di Zhang, Kun Gai

机构 * Kling Team, Kuaishou Technology(快手科技 Kling Team)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出RPE-2D,通过随机化位置编码提升扩散变换器的分辨率泛化能力,实现高分辨率和低分辨率图像生成的无缝过渡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07960 2026-01-08 cs.CV 83%

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning

VisualCloze: 一种通过视觉上下文学习实现通用图像生成的框架

Zhong-Yu Li, Ruoyi Du, Juncheng Yan, Le Zhuo, Qilong Wu, Zhen Li, Peng Gao, Zhanyu Ma, Ming-Ming Cheng

机构 * VCIP, CS, Nankai University(VCIP、计算机科学、南开大学) Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 VisualCloze通过视觉上下文学习实现通用图像生成,支持多种任务泛化与反向生成,利用图结构数据集提升任务密度和知识迁移。

Comments Accepted at ICCV 2025. Project page: https://visualcloze.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02036 2026-01-06 cs.LG cs.CV 83%

GDRO: Group-level Reward Post-training Suitable for Diffusion Models

GDRO:适用于扩散模型的组级奖励后训练

Yiyang Wang, Xi Chen, Xiaogang Xu, Yu Liu, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学) Tongyi Lab(通义实验室)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 GDRO通过组级奖励后训练优化,提升扩散模型奖励分数并增强抗奖励黑客能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23239 2026-01-06 cs.CV 83%

RS-Prune: Training-Free Data Pruning at High Ratios for Efficient Remote Sensing Diffusion Foundation Models

RS-Prune: 无需训练的数据裁剪以高比例提升高效遥感扩散基础模型

Fan Wei, Runmin Dong, Yushan Lai, Yixiang Yang, Zhaoyang Luo, Jinxiao Zhang, Miao Yang, Shuai Yuan, Jiyao Zhao, Bin Luo, Haohuan Fu

机构 * Department of Earth System Science, Tsinghua University(清华大学地球系统科学系) School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院) National Supercomputing Center in Shenzhen(深圳国家超算中心) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生学院) Department of Geography, The University of Hong Kong(香港大学地理系) Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 RS-Prune提出一种无需训练的高比例数据裁剪方法,通过局部信息与全局多样性结合,提升遥感扩散模型的收敛性和生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01487 2026-01-06 cs.CV cs.AI 83%

DeepInv: A Novel Self-supervised Learning Approach for Fast and Accurate Diffusion Inversion

DeepInv: 一种用于快速准确扩散逆向的新型自监督学习方法

Ziyue Zhang, Luxi Lin, Xiaolin Hu, Chao Chang, HuaiXi Wang, Yiyi Zhou, Rongrong Ji

机构 * Xiamen University(厦门大学) National University of Defense Technology(国防科技大学)

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV

AI总结 DeepInv通过自监督学习和数据增强策略,实现快速准确的扩散逆向,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00546 2026-01-06 cs.CV 83%

Seal2Real: Prompt Prior Learning on Diffusion Model for Unsupervised Document Seal Data Generation and Realisation

Seal2Real: 基于扩散模型的提示先验学习用于无监督文档印章数据生成与实现

Mingfu Yan, Jiancheng Huang, Shifeng Chen

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Shenzhen University of Advanced Technology(深圳大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 Seal2Real通过基于预训练扩散模型的提示先验学习,生成大规模标注的文档印章数据,提升真实数据中印章相关任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00535 2026-01-05 cs.CV 83%

FreeText: Training-Free Text Rendering in Diffusion Transformers via Attention Localization and Spectral Glyph Injection

FreeText: 通过注意力定位和频谱字形注入实现免训练的文本渲染

Ruiqiang Zhang, Hengyi Wang, Chang Liu, Guanjie Wang, Zehua Ma, Weiming Zhang

机构 * Anhui Province Key Laboratory of Digital Security, University of Science and Technology of China(安徽省数字安全重点实验室,中国科学技术大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 FreeText通过注意力定位和频谱字形注入技术,无需训练即可提升文本渲染的准确性和美学质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24195 2026-01-01 cs.CV 83%

CorGi: Contribution-Guided Block-Wise Interval Caching for Training-Free Acceleration of Diffusion Transformers

CorGi:基于贡献的块级区间缓存用于无训练加速扩散变换器

Yonglak Son, Suhyeok Kim, Seungryong Kim, Young Geun Kim

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 CorGi通过块级区间缓存减少DiT去噪过程中的冗余计算,提升推理速度并保持生成质量,CorGi+进一步优化了文本到图像任务中的注意力更新。

Comments 16 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24138 2026-01-01 cs.LG cs.AI cs.CV 83%

GARDO: Reinforcing Diffusion Models without Reward Hacking

GARDO:无需奖励黑客的扩散模型强化

Haoran He, Yuxiao Ye, Jie Liu, Jiajun Liang, Zhiyong Wang, Ziyang Yuan, Xintao Wang, Hangyu Mao, Pengfei Wan, Ling Pan

机构 * Hong Kong University of Science and Technology(香港科学与技术大学) Kuaishou Technology(快手科技) CUHK MMLab(港中文大学MMLab) The University of Edinburgh(爱丁堡大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 GARDO通过自适应正则化和多样性增强,有效缓解扩散模型中的奖励黑客问题,提升生成多样性与样本效率。

Comments 17 pages. Project: https://tinnerhrhe.github.io/gardo_project

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15560 2025-12-30 cs.CV 83%

GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models

GRAN-TED: 生成稳健、对齐且细腻的文本嵌入以用于扩散模型

Bozhou Li, Sihan Yang, Yushuo Guan, Ruichuan An, Xinlong Chen, Yang Shi, Pengfei Wan, Wentao Zhang, Yuanxing zhang

机构 * Peking University(北京大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Xi’an Jiaotong University(西安交通大学) School of Artificial Intelligence, UCAS(清华大学人工智能学院)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 GRAN-TED提出了一种生成稳健、对齐且细腻的文本嵌入以用于扩散模型的范式,通过TED-6K基准提升文本编码器性能,显著加快扩散模型训练速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17951 2025-12-30 cs.CV cs.AI 83%

Robust Polyp Detection and Diagnosis through Compositional Prompt-Guided Diffusion Models

通过组成提示引导的扩散模型实现鲁棒的息肉检测与诊断

Jia Yu, Yan Zhu, Peiyao Fu, Tianyi Chen, Junbo Huang, Quanlin Li, Pinghong Zhou, Zhihua Wang, Fei Wu, Shuo Wang, Xian Yang

机构 * Zhejiang University(浙江大学) Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海高等研究院) Endoscopy Center, Zhongshan Hospital, Fudan University(复旦大学中山医院内窥镜中心) Shanghai Collaborative Innovation Center of Endoscopy(上海内窥镜协同创新中心) Digital Medical Research Center, School of Basic Medical Sciences, Fudan University(复旦大学基础医学院数字医疗研究中心) Shanghai Key Laboratory of MICCAI(上海MICCAI重点实验室) Alliance Manchester Business School, The University of Manchester(曼彻斯特大学联盟商学院) Data Science Institute, Imperial College London(伦敦帝国学院数据科学研究所)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出PSDM模型,通过整合多种临床注释生成合成图像,提升息肉检测、分类和分割的鲁棒性和泛化能力。

Journal ref IEEE Trans. Med. Imaging, vol. 44, no. 12, pp. 5245-5257, Dec. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22626 2025-12-30 cs.CV 83%

Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion

Envision:通过目标图像视频扩散进行具身视觉规划

Yuming Gu, Yizhi Wang, Yining Hong, Yipeng Gao, Hao Jiang, Angtian Wang, Bo Liu, Nathaniel S. Dennler, Zhengfei Kuang, Hao Li, Gordon Wetzstein, Chongyang Ma

机构 * University of Southern California(南加州大学) ByteDance(字节跳动) Stanford University(斯坦福大学) Massachusetts Institute of Technology(麻省理工学院) MBZUAI

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV

AI总结 Envision 通过目标图像视频扩散模型实现具身视觉规划,提升目标对齐和空间一致性,支持下游机器人任务。

Comments Page: https://envision-paper.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22323 2025-12-30 cs.CV cs.AI 83%

SpotEdit: Selective Region Editing in Diffusion Transformers

SpotEdit: 差分变换器中的选择性区域编辑

Zhibin Qin, Zhenxiong Tan, Zeqing Wang, Songhua Liu, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV

AI总结 SpotEdit提出了一种无需训练的差分编辑框架,通过选择性更新修改区域,实现高效且精确的图像编辑。

Comments 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21593 2025-12-29 stat.ML cs.AI cs.CV cs.LG 83%

Residual Prior Diffusion: A Probabilistic Framework Integrating Coarse Latent Priors with Diffusion Models

残差先验扩散:整合粗粒度潜在先验与扩散模型的概率框架

Takuro Kutsuna

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 残差先验扩散通过整合粗粒度潜在先验与扩散模型,有效捕捉数据分布的细粒度细节并保持大型结构,提升生成质量。

Comments 40 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20233 2025-12-24 cs.LG cs.CV 83%

How I Met Your Bias: Investigating Bias Amplification in Diffusion Models

如何遇见你的偏见:调查扩散模型中的偏见放大

Nathan Roos, Ekaterina Iakovleva, Ani Gjergji, Vito Paolo Pastore, Enzo Tartaglione

机构 * LTCI, Télécom Paris, Institut Polytechnique de Paris(LTCI、 Télécom Paris、 Institut Polytechnique de Paris) MaLGa-DIBRIS, University of Genova(MaLGa-DIBRIS、乌尔比诺大学) AIGO, Istituto Italiano di Tecnologia(AIGO、意大利技术研究所)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 本研究探讨扩散模型中采样算法和超参数对偏见放大影响,通过实验展示超参数可调节偏见程度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19479 2025-12-23 cs.CV 83%

Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation

情绪导演:在情绪导向图像生成中弥合情感捷径

Guoli Jia, Junyao Hu, Xinwei Long, Kai Tian, Kaiyan Zhang, KaiKai Zhao, Ning Ding, Bowen Zhou

机构 * Tsinghua University(清华大学) The Hong Kong Polytechnic University(香港理工大学) China Unicom(中国联合通信) Shanghai Artificial Intelligence Lab(上海人工智能实验室)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 Emotion-Director通过跨模态协作框架,结合MC-Diffusion和MC-Agent,提升情绪导向图像生成的准确性和表现力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18254 2025-12-23 cs.CV 83%

Loom: Diffusion-Transformer for Interleaved Generation

Loom:用于交错生成的扩散-变压器

Mingcheng Ye, Jiaming Liu, Yiren Song

机构 * Beijing Institute of Technology(北京理工大学) National University of Singapore(新加坡国立大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 Loom通过统一的扩散-变压器框架实现交错生成,提供更高效和可控的长周期生成,优于现有基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18161 2025-12-23 cs.CV 83%

Local Patches Meet Global Context: Scalable 3D Diffusion Priors for Computed Tomography Reconstruction

局部片段与全局上下文:用于计算机断层扫描重建的可扩展3D扩散先验

Taewon Yang, Jason Hu, Jeffrey A. Fessler, Liyue Shen

机构 * EECS Department, University of Michigan(密歇根大学电子工程与计算机科学系)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出一种3D片段基扩散模型,通过学习3D片段先验实现高效高分辨率3D CT重建,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17782 2025-12-22 cs.CV 83%

UrbanDIFF: A Denoising Diffusion Model for Spatial Gap Filling of Urban Land Surface Temperature Under Dense Cloud Cover

UrbanDIFF: 一种用于密集云覆盖下城市地表温度空间填补的去噪扩散模型

Arya Chavoshi, Hassan Dashtian, Naveen Sudharsan, Dev Niyogi

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 UrbanDIFF是一种基于去噪扩散模型的城市地表温度空间填补方法,通过静态城市结构信息和监督像素引导细化步骤,有效应对密集云覆盖下的LST重建问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17578 2025-12-22 cs.CV 83%

3One2: One-step Regression Plus One-step Diffusion for One-hot Modulation in Dual-path Video Snapshot Compressive Imaging

3One2: 一步回归加一步扩散用于双路径视频压缩成像中的one-hot调制

Ge Wang, Xing Liu, Xin Yuan

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 3One2通过结合一步回归和一步扩散方法,解决双路径视频压缩成像中的one-hot调制问题,实现高效的时间解耦和空间增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17573 2025-12-22 cs.CV 83%

RoomEditor++: A Parameter-Sharing Diffusion Architecture for High-Fidelity Furniture Synthesis

RoomEditor++: 一种用于高保真家具合成的参数共享扩散架构

Qilong Wang, Xiaofan Ming, Zhenyi Lin, Jinwen Li, Dongwei Ren, Wangmeng Zuo, Qinghua Hu

机构 * School of Artificial Intelligence, Tianjin University(天津大学人工智能学院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) Engineering Research Center of City Intelligence and Digital Governance, Ministry of Education(教育部城市智能与数字治理工程研究中心)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 RoomEditor++通过参数共享的扩散架构实现高保真家具合成,提供公开数据集并优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07449 2025-12-22 cs.CV 83%

Tracing the Roots: Leveraging Temporal Dynamics in Diffusion Trajectories for Origin Attribution

追溯根源:利用扩散轨迹的时间动态性进行起源归因

Andreas Floros, Seyed-Mohsen Moosavi-Dezfooli, Pier Luigi Dragotti

机构 * Imperial College London(伦敦帝国学院)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 本文提出利用扩散轨迹的时间动态性,解决图像起源归因问题,挑战现有成员推断方法,并提出统一的数据来源框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16586 2025-12-19 cs.CV cs.AI 83%

Yuan-TecSwin: A text conditioned Diffusion model with Swin-transformer blocks

Yuan-TecSwin: 一种基于Swin-transformer块的文本条件扩散模型

Shaohua Wu, Tong Yu, Shenling Wang, Xudong Zhao

机构 * Yuan-Tec(元 Tec)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 Yuan-TecSwin通过引入Swin-transformer块替代CNN,提升文本条件扩散模型在图像生成中的非局部建模能力,达到ImageNet生成基准的最优FID分数1.37。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05005 2025-12-18 cs.CV cs.LG cs.RO 83%

Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models

Diff-2-in-1:通过扩散模型弥合生成与密集感知的鸿沟

Shuhong Zheng, Zhipeng Bao, Ruoyu Zhao, Martial Hebert, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Carnegie Mellon University(卡内基梅隆大学) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 Diff-2-in-1通过结合扩散模型的生成与感知能力,实现多模态数据生成与密集视觉感知的统一框架,提升视觉感知的判别能力。

Comments 26 pages, 14 figures

Journal ref ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07345 2025-12-17 cs.CV 83%

Debiasing Diffusion Priors via 3D Attention for Consistent Gaussian Splatting

通过3D注意力解偏扩散先验以实现一致的高斯点散布

Shilong Jin, Haoran Duan, Litao Hua, Wentao Huang, Yuan Zhou

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 通过3D注意力解偏扩散先验以实现一致的高斯点散布,解决T2I模型中的视角偏见问题,提升3D任务的多视角一致性。

Comments Accepted by AAAI 2026, Code is available at: https://github.com/kimslong/AAAI26-TDAttn

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13991 2025-12-17 cs.CV 83%

Repurposing 2D Diffusion Models for 3D Shape Completion

将2D扩散模型用于3D形状补全

Yao He, Youngjoong Kwon, Tiange Xiang, Wenxiao Cai, Ehsan Adeli

机构 * Stanford University(斯坦福大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 本文提出利用2D扩散模型进行3D形状补全,通过Shape Atlas实现模态对齐,提升生成效果并验证其在点云补全和网格生成中的实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11203 2025-12-16 cs.CV 83%

AutoRefiner: Improving Autoregressive Video Diffusion Models via Reflective Refinement Over the Stochastic Sampling Path

AutoRefiner: 通过在随机采样路径上的反思细化改进自回归视频扩散模型

Zhengyang Yu, Akio Hayakawa, Masato Ishii, Qingtao Yu, Takashi Shibuya, Jing Zhang, Yuki Mitsufuji

机构 * Australian National University(澳大利亚国立大学) Sony AI(索尼人工智能) Sony Group Corporation(索尼集团)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 AutoRefiner通过路径噪声细化和反射式KV缓存改进AR-VDMs的样本保真度

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08987 2025-12-16 cs.CV 83%

WDT-MD: Wavelet Diffusion Transformers for Microaneurysm Detection in Fundus Images

WDT-MD:用于视网膜图像微动脉瘤检测的小波扩散变换器

Yifei Sun, Yuzhi He, Junhao Jia, Jinhong Wang, Ruiquan Ge, Changmiao Wang, Hongxia Xu

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 WDT-MD 通过小波扩散变换器框架,解决微动脉瘤检测中的身份映射、区分困难和重建效果差问题,提升视网膜图像筛查性能。

Comments 9 pages, 6 figures, 8 tables, accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13080 2025-12-16 cs.CV 83%

Counting Hallucinations in Diffusion Models

扩散模型中hallucination的计数

Shuai Fu, Jian Zhou, Qi Chen, Huang Jing, Huy Anh Nguyen, Xiaohan Liu, Zhixiong Zeng, Lin Ma, Quanshi Zhang, Qi Wu

机构 * University of Adelaide(阿德莱德大学) Meituan(美团) Shanghai Jiao Tong University(上海交通大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出了一种针对扩散模型中计数hallucination的评估方法,通过构建CountHalluSet数据集,系统分析不同采样条件对hallucination的影响,并揭示FID等指标无法捕捉计数hallucination的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10041 2025-12-15 cs.CV cs.AI 83%

MetaVoxel: Joint Diffusion Modeling of Imaging and Clinical Metadata

MetaVoxel:联合图像与临床元数据的扩散建模

Yihao Liu, Chenyu Gao, Lianrui Zuo, Michael E. Kim, Brian D. Boyd, Lisa L. Barnes, Walter A. Kukull, Lori L. Beason-Held, Susan M. Resnick, Timothy J. Hohman, Warren D. Taylor, Bennett A. Landman

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 MetaVoxel通过联合扩散建模统一图像与临床元数据,实现图像生成、年龄估计和性别预测,性能媲美传统任务特定模型。

详情

展开后加载摘要…

URL PDF HTML 收藏