arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

European Conference on Computer Vision · 会议 · Computer Vision

共收录 4781
2607.01906 2026-07-03 cs.CV 新提交

SFKD: Spatial--Frequency Joint-Aware Heterogeneous Knowledge Distillation via Multi-Level Wavelet Spectral Interaction

SFKD: 通过多级小波频谱交互的空间-频率联合感知异质知识蒸馏

Cuipeng Wang, Haipeng Wang

机构 * Key Laboratory for Information Science of Electromagnetic Waves, Ministry of Education, Fudan University(复旦大学电磁波信息科学教育部重点实验室) Discipline and Technology Center of Microwave Vision Intelligent Sensing, Fudan University(复旦大学微波视觉智能感知学科与技术中心)

AI总结 提出空间-频率联合感知异质知识蒸馏框架SFKD,利用多级离散小波变换解耦空间信息,结合双流双阶段精化模块和高斯滤波频率损失,实现异质模型间的有效知识迁移。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01869 2026-07-03 cs.CV 新提交

QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers

QWERTY: 通过查询扭曲的视频扩散变换器实现无需训练的运动控制

Kyobin Choo, Youngmin Kim, Hyunkyung Han, Geunrip Park, Chanyoung Kim, Sunyoung Jung, Seong Jae Hwang

机构 * Department of Computer Science, Yonsei University(延世大学计算机科学系) Department of Artificial Intelligence, Yonsei University(延世大学人工智能系)

AI总结 提出QWERTY框架,通过扭曲查询的语义子空间,在预训练图像到视频扩散变换器中实现无需训练的运动控制,性能接近微调方法。

Comments 37 pages, 18 figures, accepted at the European Conference on Computer Vision (ECCV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01851 2026-07-03 cs.CV 新提交

Geometric Foundation Model Distillation for Efficient Lunar 3D Reconstruction

几何基础模型蒸馏用于高效月球三维重建

Clémentine Grethen, Florient Chouteau, Géraldine Morin, Simone Gasparini

机构 * IRIT, University of Toulouse(图卢兹大学IRIT研究所) Airbus Defence and Space(空中客车防务与航天公司)

AI总结 针对硬件受限场景,通过知识蒸馏将MASt3R模型压缩7倍,保留大部分重建精度,提出SVD初始化方法提升训练稳定性,并揭示编码器类型、蒸馏策略等关键因素。

Comments Accepted to ECCV 2026, code can be accessed via https://clementinegrethen.github.io/publications/ECCV.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01827 2026-07-03 cs.CV 新提交

C2E: Boosting Ego-Only 3D Object Detection via Multi-Teacher Contrastive Knowledge Distillation

C2E: 通过多教师对比知识蒸馏提升仅自我3D目标检测

Jinlong Wang, Xun Huang, Qiming Xia, Shijia Zhao, Chenglu Wen

机构 * Fujian Key Laboratory of Urban Intelligent Sensing and Computing, Xiamen University(福建省城市智能感知与计算重点实验室,厦门大学) Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(多媒体可信感知与高效计算教育部重点实验室,厦门大学) Zhongguancun Academy(中关村学院)

AI总结 提出C2E范式,通过多对一智能体对比知识蒸馏框架M2S,结合多级特征增强、辅助点云重建和多教师对比蒸馏,在无通信成本下提升仅自我3D检测性能,在多个数据集上验证有效性。

Comments 18 pages, 8figures

Journal ref ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01754 2026-07-03 cs.AI 新提交

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

路径级事后指令用于视觉语言导航中的语义探索

Sung June Kim, Sangpil Kim, Honglak Lee

机构 * Korea University(高丽大学) University of Michigan(密歇根大学)

AI总结 提出Phi-Nav框架,通过事后推理将指令与智能体实际探索轨迹对齐,利用三阶段双监督循环将语义未标记的运动转化为密集训练信号,在R2R-CE和RxR-CE基准上以少量专家演示取得竞争性能。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01743 2026-07-03 cs.CV 新提交

InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation

InterCMDM:用于自回归人体交互生成的块因果扩散

Qing Yu, Kent Fujiwara

机构 * LY Corporation(LY公司)

AI总结 提出InterCMDM块因果潜扩散框架,通过双流因果扩散Transformer和多重任务注意力掩码实现自回归双人交互生成,解决因果性缺失和时间漂移问题,在InterHuman和InterX上达到最优性能。

Comments Accepted to ECCV 2026, Project website: https://yu1ut.com/InterCMDM-HP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01737 2026-07-03 cs.CV 新提交

ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA

ReQuest:基于重新思考的问题感知帧选择用于长视频问答

Minkuk Kim, Suyong Yun, Young Tae Kim, Jinyoung Moon, Jinwoo Choi, Seong Tae Kim

AI总结 提出ReQuest,一种不确定性驱动的、问题自适应的关键帧选择方法,通过轻量级选择器、重新思考路由和不确定性引导的NMS,在固定token预算下提升长视频问答准确率。

Comments Accepted at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01707 2026-07-03 cs.CV 新提交

LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression

LASER: 通过视觉注意力保持与沉没抑制为LVLMs提供矫正透镜

Bowen Yuan, Zijian Wang, Yadan Luo, Shijie Wang, Zi Huang

机构 * The University of Queensland(昆士兰大学)

AI总结 针对大型视觉语言模型在长序列解码中视觉遗忘的问题,提出LASER后训练框架,通过视觉接地奖励和沉没抑制奖励调节注意力轨迹与分布,在8个基准上超越强基线。

Comments The 19th European Conference on Computer Vision (ECCV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01677 2026-07-03 cs.CV 新提交

ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning

ICDepth: 通过上下文条件化驯服视频扩散模型用于视频深度估计

Xuanhua He, Jiaxin Xie, Mingzhe Zheng, Qifeng Chen

机构 * The Hong Kong University of Science and Technology(香港科技大学)

AI总结 提出ICDepth框架,利用上下文条件化(ICC)将预训练文本到视频扩散Transformer适配到视频深度估计,通过SAND-Attention确保时空对齐和SRFM注入语义分辨率先验,仅用0.8M帧训练数据即达到SOTA并展现强零样本泛化。

Comments Accepted to ECCV 2026. Project page: https://xuanhuahe.github.io/ICDepth/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01667 2026-07-03 cs.CV 新提交

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning

增强视听视频字幕的时间与跨模态对齐

Chen Zhao, Jiajun Ma, Qilong Huang, Tiehan Fan, Hongyu Li, Zhuoliang Kang, Xiaoming Wei, Jian Yang, Ying Tai

机构 * Nanjing University, State Key Laboratory for Novel Software Technology(南京大学,计算机软件新技术国家重点实验室) Meituan(美团)

AI总结 提出TCA-Captioner框架,通过观察-检查-纠正迭代策略和密集交互数据集,解决视听字幕中的模态分离与时间不一致问题,实现精准的视听绑定与因果动态建模。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01663 2026-07-03 cs.CV 新提交

Unified Panoramic-Gaussian Representation for Monocular 4D Scene Synthesis

统一全景-高斯表示用于单目4D场景合成

Yuankun Yang, Yi Wei, Wenyang Zhou, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院) Central Media Technology Institute, Huawei(华为中央媒体技术研究院)

AI总结 提出PanoGaussian,结合全景轨迹引导与显式动态高斯表示,解决单目视频4D场景合成中视角外区域推断不一致和动态内容建模问题。

Comments Accepted at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01658 2026-07-03 cs.CV 新提交

Teaching Vision-Language-Action Models What to See and Where to Look

教视觉-语言-动作模型看什么和看哪里

Yuguang Yang, Canyu Chen, Zhewen Tan, Yizhi Wang, Zichao Feng, Chunyang Liu, Kehua Sheng, Juan Zhang, Linlin Yang, Baochang Zhang, Yan Wang, Bo Zhang, Xianbin Cao

机构 * School of Electronic Information Engineering, Beihang University(北京航空航天大学电子信息工程学院) National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) DiDi(滴滴出行) State Key Laboratory of Media Convergence and Communication, Communication University of China(中国传媒大学媒体融合与传播国家重点实验室) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) School of Cyber Science and Technology, Beihang University(北京航空航天大学网络安全科学与技术学院) School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)

AI总结 提出DriveTeach-VLA框架,通过驾驶感知视觉蒸馏和2D轨迹引导提示,教VLA模型关注驾驶相关区域,实现端到端自动驾驶轨迹预测,在NAVSIM和nuScenes上达到最优性能。

Comments The paper has been accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01657 2026-07-03 cs.CV 新提交

Domain Generalization via Text-Anchored Information Bottleneck

基于文本锚定信息瓶颈的域泛化

Eunyi Lyou, Yunjeong Choi, Junho Lee, Joonseok Lee

机构 * Seoul National University(首尔大学)

AI总结 提出通过语言嵌入空间作为信息瓶颈来抑制视觉表征中的虚假线索,从而提升域泛化性能,实验表明该方法达到最新最优水平。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01503 2026-07-03 cs.CV 新提交

Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task

通过深度排序任务解构视觉语言模型中的图像线索理解与语言偏差

Yiqian Liu, Iuliia Kotseruba, John K. Tsotsos

机构 * York University(约克大学) University of Guelph(圭尔夫大学)

AI总结 提出深度排序与异常检测任务,结合控制图像线索和语言表达,量化视觉语言模型的深度感知能力,发现模型对深度线索利用不足且存在语言偏差。

Comments 15 pages, 7 figures, accepted to ECCV 2026 (30 pages, 13 figures, supplementary materials included)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01370 2026-07-03 cs.CV 新提交

MapDreamer: Aerial Imagery Conditioned Latent Diffusion for Lane-Level Map Generation

MapDreamer: 基于航空影像条件潜扩散的车道级地图生成

Julian Brandes, Philipp Crocoll, Wolfram Burgard

机构 * Department CSAI at University of Technology Nuremberg(纽伦堡工业大学CSAI系) Robert Bosch GmbH(罗伯特·博世有限公司)

AI总结 提出MapDreamer,一种从单张航空影像直接生成带显式拓扑的车道级矢量地图的生成扩散模型,通过变分自编码器学习车道中心线及其拓扑的紧凑潜表示,并利用基于Transformer的潜扩散模型预测图结构,在几何和拓扑保真度上优于非生成基线。

Comments Accepted at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01272 2026-07-03 cs.GR cs.AI cs.CV cs.DC cs.LG 新提交

Benchmarking Federated Learning and Knowledge Distillation for Point Cloud Classification

联邦学习与知识蒸馏在点云分类中的基准测试

Aizierjiang Aiersilan

机构 * University of Macau(澳门大学)

AI总结 针对隐私敏感和资源受限场景,联合评估联邦学习与知识蒸馏在3D点云分类中的性能,发现极端非独立同分布标签偏移下联邦学习性能下降,蒸馏可压缩模型且避免标签泄露问题。

Comments We are pleased to announce that this paper has been accepted by the 19th European Conference on Computer Vision (ECCV 2026). We appreciate the valuable feedback from the reviewers and look forward to sharing our findings with the community

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28864 2026-07-03 cs.CV 版本更新

On Test-Time Scaling for Vision-Language Models

关于视觉-语言模型的测试时扩展

Fawaz Sammani, Tzoulio Chamiti, Nikos Deligiannis

机构 * ETRO Department, Vrije Universiteit Brussel(布鲁塞尔自由大学ETRO部门) imec

AI总结 本文系统研究了视觉-语言模型(LVLM)的测试时扩展方法,发现小型高性能模型受益最大,性能提升可达30%,并揭示了视觉信息在推理链早期编码后图像令牌贡献显著下降。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31918 2026-07-03 cs.CV 新提交

DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation

DriveWeaver: 基于点云条件的视频修复用于自动驾驶仿真中的可控车辆插入

Junzhe Jiang, Zipei Ma, Zijie Pan, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院) Shanghai Innovation Institute(上海创新研究院)

AI总结 提出DriveWeaver框架,通过点云条件视频修复实现自动驾驶仿真中可控车辆插入,解决光照不一致和3D资产依赖问题,支持大规模场景增强。

Comments Accepted at ECCV 2026, Project Page: https://github.com/LogosRoboticsGroup/DriveWeaver

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31699 2026-07-03 cs.CV cs.AI 新提交

Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models

用稀疏自编码器实现扩散模型中的遗忘:只看不碰

Enrico Cassano, Riccardo Renzulli, Rayyan Ahmed, Marco Grangetto, Stephan Alaniz

机构 * University of Turin(都灵大学) LTCI, Télécom Paris, Institut Polytechnique de Paris(LTCI,巴黎电信学院,巴黎综合理工学院)

AI总结 本文利用稀疏自编码器作为语义检测器,通过检测并替换包含目标对象的图像区域,实现扩散模型中的对象擦除,避免了直接干预潜在空间导致的视觉伪影。

Comments Accepted at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29691 2026-07-03 cs.CV 版本更新

Unsupervised Semantic Segmentation Facilitates Model Understanding

无监督语义分割促进模型理解

Xiaoyan Yu, Lisa Mais, Jannik Franzen, Peter Hirsch, Nick Lechtenbörger, Andreas Mardt, Dagmar Kainmüller

机构 * Max-Delbruck-Center(马克斯·德尔布鲁克中心) Helmholtz Imaging(海德堡成像) Humboldt-Universität zu Berlin(柏林洪堡大学) Charité Universitätsmedizin(夏里特大学医学院) University of Potsdam(波茨坦大学)

AI总结 提出基于无监督语义分割的可视化协议,直观揭示不同自监督视觉Transformer的注意力机制、位置偏差和缩放行为等模型特性。

Comments Camera-ready version of paper accepted at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26283 2026-07-03 cs.CV cs.AI 版本更新

MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution

MedSynapse-V:通过潜在记忆演化桥接视觉感知与临床直觉

Chunzheng Zhu, Jiaqi Zeng, Junyu Jiang, Jianxin Lin, Yijun Wang

机构 * Hunan University(湖南大学)

AI总结 提出MedSynapse-V框架,通过潜在诊断记忆演化模拟临床专家经验调用,解决医学视觉语言模型因离散分词导致的量化损失、长程信息消散和案例适应性问题,在诊断准确性上显著超越现有方法。

Comments ECCV 2026; Medical latent reasoning; Memory evolution

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02546 2026-07-03 cs.CV cs.LG 版本更新

RGB-Pointmap Pretraining for Unified 3D Scene Understanding

对比语言-彩色点图预训练用于统一3D场景理解

Ye Mao, Weixun Luo, Ranran Huang, Junpeng Jing, Krystian Mikolajczyk

机构 * Imperial College London(帝国理工学院伦敦分校)

AI总结 提出UniScene3D,一种基于Transformer的编码器,通过多视图彩色点图联合建模外观和几何,并引入跨视图几何对齐和接地视图对齐,在少样本和任务特定微调中达到SOTA。

Comments 19 Pages, ECCV 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01761 2026-07-03 cs.CV 版本更新

Control-DINO: Feature Space Conditioning for Controllable Image-to-Video Diffusion

Control-DINO:用于可控图像到视频扩散的特征空间条件

Edoardo A. Dominici, Thomas Deixelberger, Konstantinos Vardis, Markus Steinberger

机构 * Huawei Technologies, Switzerland(华为技术(瑞士)) Huawei Technologies, Austria(华为技术(奥地利)) Graz University of Technology, Austria(格拉茨技术大学)

AI总结 本文提出Control-DINO,通过解耦外观与其他特征,实现对视频扩散模型的可控生成,提升生成渲染的可控性。

Comments ECCV 2026 - Project Page https://dedoardo.github.io/projects/control-dino/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27999 2026-07-03 cs.CV 版本更新

CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition

CLIP-AUTT: 基于动作单元提示的测试时个性化方法用于细粒度视频情绪识别

Muhammad Osama Zeeshan, Masoumeh Sharafi, Benoit Savary, Alessandro Lameiras Koerich, Marco Pedersoli, Eric Granger

机构 * LIVIA, ILLS, Dept. of Systems Engineering, ETS Montreal(LIVIA, ILLS, 系统工程学院, 蒙特利尔高等技术学院) LIVIA, Dept. of Software and IT Engineering, ETS Montreal(LIVIA, 软件与IT工程学院, 蒙特利尔高等技术学院) Dept. of Computer Science, École Polytechnique(计算机科学系, 巴黎综合理工学院)

AI总结 本文提出CLIP-AUTT,通过动作单元提示实现测试时个性化,提升细粒度视频情绪识别的鲁棒性和个性化能力。

Comments ECCV, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28896 2026-07-03 cs.CV 版本更新

Fisheye3R: Adapting Unified 3D Feed-Forward Foundation Models to Fisheye Lenses

Fisheye3R:将统一的3D前馈基础模型适应于鱼眼镜头

Ruxiao Duan, Erin Hong, Dongxu Zhao, Eric Turner, Alex Wong, Yunwen Zhou

机构 * Yale University(耶鲁大学) Google XR(谷歌XR)

AI总结 本文提出Fisheye3R框架,通过灵活学习方案在不退化性能的前提下,使多视角3D重建模型适应高径向畸变的鱼眼图像。

Comments European Conference on Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19013 2026-07-03 cs.CV 版本更新

GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness

GenHOI:具有遮挡意识的通用手-物体姿态估计

Hui Yang, Wei Sun, Jian Liu, Jian Xiao, Tao Xie, Hossein Rahmani, Ajmal Saeed Mian, Nicu Sebe, Gim Hee Lee

机构 * Hunan University(湖南大学) Lancaster University(兰卡斯特大学) University of Western Australia(西澳大学) University of Trento(特伦托大学) National University of Singapore(新加坡国立大学)

AI总结 本文提出GenHOI框架,通过整合层次语义知识与手部先验,提升在遮挡条件下手-物体姿态估计的泛化能力,实验表明在DexYCB和HO3Dv2基准上达到SOTA性能。

Comments European Conference on Computer Vision (ECCV), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24969 2026-07-03 cs.CV 版本更新

PASDiff: Physics-Aware Semantic Guidance for Joint Real-World Low-Light Face Enhancement and Restoration

PASDiff: 面向真实世界低光人脸增强与恢复的物理感知语义引导

Yilin Ni, Wenjie Li, Zhengxue Wang, Juncheng Li, Guangwei Gao, Jian Yang

机构 * Nanjing University of Posts and Telecommunications(南京邮电大学) Beijing University of Posts and Telecommunications(北京邮电大学) PCA Lab, Nanjing University of Science and Technology(南京理工大学PCA实验室) East China Normal University(华东师范大学)

AI总结 提出PASDiff,一种无需训练的物理感知语义扩散模型,通过逆强度加权和Retinex理论引入光度约束,并利用风格无关结构注入从现成人脸先验中提取结构,在真实低光人脸增强上实现自然光照、色彩恢复和身份一致性的优越平衡。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09446 2026-07-03 cs.CV 版本更新

Defect-aware Hybrid Prompt Optimization via Progressive Tuning for Zero-Shot Multi-type Anomaly Detection and Segmentation

基于渐进调优的缺陷感知混合提示优化用于零样本多类型异常检测与分割

Nadeem Nazer, Hongkuan Zhou, Lavdim Halilaj, Ylli Sadikaj, Steffen Staab

机构 * Corporate Research, Robert Bosch GmbH(罗伯特·博世集团企业研究部) Otto-von-Guericke-University(奥托·冯·格里克大学) Institute for Artificial Intelligence, University of Stuttgart(斯图加特大学人工智能研究所) Faculty of Computer Science, UniVie Doctoral School Computer Science, University of Vienna(维也纳大学计算机科学系) University of Southampton(南安普顿大学)

AI总结 提出DAPO框架,通过混合提示机制结合可读缺陷描述与可学习嵌入,对齐异常视觉特征与文本语义,在零样本多类型异常检测与分割任务中平均提升AUROC 3.6%和5.2%。

Journal ref European Conference on Computer Vision (ECCV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08337 2026-07-03 cs.CV 版本更新

Language-Guided Transformer Tokenizer for Human Motion Generation

语言引导的Transformer分词器用于人体运动生成

Sheng Yan, Yong Wang, Xin Du, Junsong Yuan, Mengyuan Liu

机构 * Transsion Ltd.(Transsion有限公司) Chongqing University of Technology(重庆理工大学) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) Shenzhen Graduate School(深圳研究生院)

AI总结 提出语言引导分词器LG-Tok,在分词阶段对齐语言与运动,产生紧凑高层语义表示,结合Transformer注意力机制实现高效运动生成,在HumanML3D和Motion-X上达到SOTA。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12770 2026-07-03 cs.CV 版本更新

One-Shot Feed-Forward 360$^{\circ}$ Animatable Avatar via Inpainted UV-Space Gaussian Modeling

通过修补UV空间高斯建模的一次前馈360度可动画化头像

Shuling Zhao, Dan Xu

机构 * The Hong Kong University of Science and Technology(香港科技大学) Zeekr Automobile R&D Co., Ltd(Zeekr汽车研发有限公司)

AI总结 提出一种基于UV空间高斯建模的一次前馈框架,利用预训练3D GAN的先验知识修补缺失区域,实现360度全头可动画化头像的高质量重建与实时动画。

Comments Accepted by ECCV 2026. Project page: https://shaelynz.github.io/fhavatar/

详情

展开后加载摘要…

URL PDF HTML 收藏