arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-02-03 至 2026-02-03 共收录 6 信号源:cs.CV, cs.GR, cs.MM

1. 效率与蒸馏 6 篇

2512.11831 2026-02-03 cs.LG cs.CV 79%

On the Design of One-step Diffusion via Shortcutting Flow Paths

关于通过捷径流路径设计一步扩散模型

Haitao Lin, Peiyan Hu, Minsi Ren, Zhifeng Gao, Zhi-Ming Ma, Guolin ke, Tailin Wu, Stan Z. Li

机构 * Department of Artificial Intelligence, School of Engineering, Westlake University(人工智能系,工程学院,西湖大学) Academy of Mathematics and Systems Science, Chinese Academy of Sciences(数学与系统科学学院,中国科学院) DP Technology, Beijing(北京DP技术)

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种通用设计框架,用于改进捷径模型,使一步生成模型在无分类器指导设置下达到新的SOTA FID50k值。

Comments 10 pages of main body, conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02390 2026-02-03 cs.GR cs.AI eess.IV 79%

F-scheduler: illuminating the free-lunch design space for fast sampling of diffusion models

F-scheduler: 探索扩散模型快速采样中的免费午餐设计空间

Zilai Li, Lujia Bai

机构 * Department of Mathematics, at Ruhr University Bochum in the group of Holger Dette, Germany(鲁尔大学博德姆数学系) Independent Researcher, Guangdong, China(独立研究者)

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.GR

AI总结 F-scheduler通过优化ODE求解器与Free-U Net结合,实现扩散模型快速采样,提升图像生成质量与效率

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00107 2026-02-03 cs.CV cs.RO eess.IV 74%

Efficient UAV trajectory prediction: A multi-modal deep diffusion framework

高效无人机轨迹预测:一种多模态深度扩散框架

Yuan Gao, Xinyu Guo, Wenjing Xie, Zifan Wang, Hongwen Yu, Gongyang Li, Shugong Xu

专题命中 效率与蒸馏 :diffusion(title);分类 cs.CV

AI总结 本文提出一种多模态深度融合框架,通过融合激光雷达和毫米波雷达数据提升无人机轨迹预测精度,实验显示其比基线模型提升40%。

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01814 2026-02-03 cs.CV 57%

GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation

GPD: 用于快速高质量视频生成的引导渐进蒸馏

Xiao Liang, Yunzhu Zhang, Linchao Zhu

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

专题命中 效率与蒸馏 :diffusion(abstract);分类 cs.CV

AI总结 GPD通过引导渐进蒸馏方法,减少视频生成的采样步骤,同时保持高质量输出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01140 2026-02-03 cs.LG 50%

Generalized Radius and Integrated Codebook Transforms for Differentiable Vector Quantization

通用半径与集成代码本变换用于可微向量量化

Haochen You, Heng Zhang, Hongyang He, Yuqi Li, Baojing Liu

机构 * Columbia University(哥伦比亚大学) South China Normal University(南方科技大学) University of Warwick(沃里克大学) The City College of New York(纽约城市学院) Hebei Institute of Communications(河北通信学院)

专题命中 效率与蒸馏 :image generation(abstract)

AI总结 GRIT-VQ通过统一的可微框架提升向量量化性能,实现稳定梯度和高效代码本利用。

Comments This paper has been accepted as a conference paper at CPAL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00641 2026-02-03 stat.ML cs.LG stat.CO 50%

Sampling from multi-modal distributions on Riemannian manifolds with training-free stochastic interpolants

在黎曼流形上采样多模分布的无训练随机插值法

Alain Durmus, Maxence Noble, Thibaut Pellerin

机构 * CMAP, CNRS, Ecole polytechnique(CMAP、法国国家科学研究中心、巴黎高等师范学院)

专题命中 效率与蒸馏 :diffusion(abstract)

AI总结 本文提出了一种无训练的随机插值方法,用于在黎曼流形上高效采样多模分布,通过非平衡动力学和随机插值实现无需训练的高效采样。

详情

展开后加载摘要…

URL PDF HTML 收藏