arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-01-09 至 2026-01-09 共收录 2 信号源:cs.CV, cs.GR, cs.MM

1. 效率与蒸馏 2 篇

2505.21400 2026-01-09 cs.LG cs.IT math.IT math.ST stat.ML stat.TH 78%

Breaking AR's Sampling Bottleneck: Provable Acceleration via Diffusion Language Models

突破生成模型的采样瓶颈:通过扩散语言模型实现可证明的加速

Gen Li, Changxiao Cai

机构 * Department of Statistics and Data Science, Chinese University of Hong Kong, Hong Kong(统计与数据科学系,香港中文大学) Department of Industrial and Operations Engineering, University of Michigan, Ann Arbor, USA(工业与运营管理系,密歇根大学)

专题命中 效率与蒸馏 :diffusion(title,abstract)

AI总结 本文从信息论角度为扩散语言模型提供收敛保证,证明采样误差随迭代次数减少而降低,从而突破自回归模型所需的L步瓶颈,为生成高质量样本提供理论支持。

Comments This is the full version of a paper published at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04792 2026-01-09 cs.CV 57%

PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference

PyramidalWan: 关于使预训练视频模型金字塔化以实现高效推理

Denis Korzhenkov, Adil Karjauv, Animesh Karnewar, Mohsen Ghafoorian, Amirhossein Habibian

机构 * Qualcomm AI Research(高通人工智能研究)

专题命中 效率与蒸馏 :diffusion(abstract);分类 cs.CV

AI总结 PyramidalWan通过低成本微调将预训练扩散模型转换为金字塔模型,提升视频推理效率并保持输出质量。

详情

展开后加载摘要…

URL PDF HTML 收藏