arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

2026-08-11 至 2026-08-11 共收录 4
2510.20512 2026-08-11 cs.CV 版本更新

Adversarial Concept Distillation for One-Step Diffusion Personalization

对抗概念蒸馏用于一步扩散个性化

Yixiong Yang, Tao Wu, Senmao Li, Shiqi Yang, Yaxing Wang, Joost van de Weijer, Kai Wang

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Computer Vision Center(计算机视觉中心) Universitat Autònoma de Barcelona(巴塞罗那自治大学) VCIP, CS, Nankai University(南开大学计算机学院视觉计算与智能感知实验室) Program of Computer Science, City University of Hong Kong (Dongguan)(香港城市大学(东莞)计算机科学项目) City University of Hong Kong(香港城市大学)

AI总结 本文提出OPAD框架,结合教师-学生蒸馏与对抗监督,通过多步扩散模型作为教师,一步学生模型联合训练,以实现一步扩散模型的高效个性化。

Comments Accepted to CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16758 2026-08-11 cs.CV 版本更新

Motion-Aware Animatable Gaussian Avatars Deblurring

具有运动感知的可动画高斯人像去模糊

Muyao Niu, Yifan Zhan, Qingtian Zhu, Zhuoxiao Li, Wei Wang, Zhihang Zhong, Xiao Sun, Yinqiang Zheng

机构 * The University of Tokyo(东京大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

AI总结 本文提出了一种从模糊视频中直接重建清晰3D人体高斯人像的方法,结合了基于物理的模糊模型和运动模型,以解决运动引起的模糊问题。

Comments Accepted at CVPR 2026, Codes: https://github.com/MyNiuuu/MAD-Avatar

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02270 2026-08-11 cs.CV 版本更新

From Visual to Multimodal: Systematic Ablation of Encoders and Fusion Strategies in Animal Identification

从视觉到多模态:动物识别中编码器和融合策略的系统消融

Vasiliy Kudryavtsev, Kirill Borodin, German Berezin, Kirill Bubenchikov, Grach Mkrtchian, Alexander Ryzhkov

机构 * Faculty of IT, Technical University of Communication(信息科技学院,通信技术大学) AI lab, Avito(人工智能实验室,Avito)

AI总结 本研究提出多模态验证框架,通过合成文本描述提升视觉特征,利用门控融合机制在动物识别中实现84.28%的准确率,比单一模态基线提升11%。

Comments Accepted to the FGVC13 Workshop at CVPR 2026. And published at MDPI Journal of Imaging (see at https://www.mdpi.com/2313-433X/12/1/30)

Journal ref Journal of Imaging (2026) 12, no. 1: 30

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07223 2026-08-11 cs.AI 版本更新

Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response

先反射,后反思:面向动态响应的延迟感知具身大语言模型智能体

Yangqing Zheng, Shunqi Mao, Dingxin Zhang, Weidong Cai

机构 * School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)

AI总结 该研究针对动态环境中具身LLM智能体的推理延迟问题,提出RRARA智能体及相关评估指标,通过时间转换机制与预规划器实现决策质量与响应能力的平衡。

Comments Accepted by the CVPR 2025 Embodied AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏