arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Huazhong University of Science and Technology(华中科技大学)

2025-12-01 至 2025-12-01 共收录 6
2511.23469 2025-12-01 cs.CV

Visual Generation Tuning

视觉生成微调

Jiahao Guo, Sinan Du, Jingfeng Yao, Wenyu Liu, Bo Li, Haoxiang Cao, Kun Gai, Chun Yuan, Kai Wu, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Tsinghua University(清华大学) School of Artificial Intelligence, South China Normal University(华南师范大学人工智能学院) Kolors Team, Kuaishou Technology(快手科技Kolors团队)

AI总结 本文提出VGT,通过视觉生成微调提升视觉语言模型的视觉生成能力,在图像重建和生成任务中均取得优异成果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23386 2025-12-01 cs.CV

VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction

VQRAE:用于多模态理解、生成和重建的表示量化自编码器

Sinan Du, Jiahao Guo, Bo Li, Shuhao Cui, Zhengzhuo Xu, Yifu Luo, Yongxian Wei, Kun Gai, Xinggang Wang, Kai Wu, Chun Yuan

机构 * Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学) Kolors Team, Kuaishou Technology(快手科技Kolors团队)

AI总结 VQRAE通过统一的分词器生成连续语义特征和离散标记,提升多模态理解、生成和重建的性能。

Comments 19 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11164 2025-12-01 cs.CV

Reverberation: Learning the Latencies Before Forecasting Trajectories

回声:在预测轨迹之前学习延迟

Conghao Wong, Ziqian Zou, Beihao Xia, Xinge You

机构 * Huazhong University of Science and Technology(华中科技大学)

AI总结 Rev模型通过回声变换学习延迟偏好,实现可控的轨迹预测,提升预测准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01528 2025-12-01 cs.CL cs.AI q-bio.BM

Leveraging Biomolecule and Natural Language through Multi-Modal Learning: A Survey

利用生物分子和自然语言通过多模态学习:一篇综述

Qizhi Pei, Zhimeng Zhou, Kaiyuan Gao, Jinhua Zhu, Yue Wang, Zun Wang, Tao Qin, Lijun Wu, Rui Yan

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Zhejiang University(浙江大学) Shanghai Innovation Institute(上海创新研究院) Huazhong University of Science and Technology(华中科技大学) University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy(中关村学院) Shanghai AI Laboratory(上海人工智能实验室) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

AI总结 本文综述了生物分子与自然语言多模态学习的最新进展,探讨了技术表示、多模态整合方法、应用实例及未来研究方向。

Comments 2025.11.28 Updated Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22595 2025-12-01 cs.CV

AnoRefiner: Anomaly-Aware Group-Wise Refinement for Zero-Shot Industrial Anomaly Detection

AnoRefiner: 一种面向零样本工业异常检测的异常感知分组细化方法

Dayou Huang, Feng Xue, Xurui Li, Yu Zhou

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) Hubei Key Laboratory of Smart Internet Technology, Huazhong University of Science and Technology(华中科技大学智能互联网技术重点实验室) Artificial Intelligence Research Institute, Wuhan JingCe Electronic Group Co., LTD(武汉景创电子集团有限公司人工智能研究所) Department of Information Engineering and Computer Science, University of Trento(特伦托大学信息工程与计算机科学系)

AI总结 AnoRefiner通过异常感知细化器提升零样本工业异常检测的像素级精度,采用异常分数地图和渐进式分组测试时训练策略改进模型性能。

Comments 17 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07769 2025-12-01 cs.CV

Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation

Dream4D: 通过可控视频生成与神经4D重建提升I2V向时空一致4D生成

Xiaoyan Liu, Kangrui Li, Yuehao Song, Jiaxin Liu

机构 * The Chinese University of Hong Kong(香港中文大学) The Hong Kong Polytechnic University(香港理工大学) Huazhong University of Science and Technology(华中科技大学) The University of New South Wales(新南威尔士大学)

AI总结 Dream4D通过可控视频生成与神经4D重建的协同,实现高质量的时空一致4D内容生成。

Comments Project Page: https://wanderer7-sk.github.io/Dream4D.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏