arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11876
2602.19596 2026-02-24 cs.CV

Learning Mutual View Information Graph for Adaptive Adversarial Collaborative Perception

学习适应性对抗协作感知的互视信息图

Yihang Tao, Senkang Hu, Haonan An, Zhengru Fang, Hangcheng Cao, Yuguang Fang

机构 * Hong Kong JC STEM Lab of Smart City(香港JC STEM实验室) City University of Hong Kong(香港城市大学)

AI总结 本文提出MVIG攻击,通过学习不同防御系统的漏洞知识,实现自适应对抗协作感知的攻击方法,有效降低防御成功率并暴露系统安全漏洞。

Comments Accepted by CVPR'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19585 2026-02-24 cs.MM cs.AI

Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis

三子空间解耦用于多模态情感分析

Chunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu, Rong Fu, Zhongxue Gan, Chun Ouyang

机构 * Fudan University(复旦大学) Peking University(北京大学) University of Macau(澳门大学)

AI总结 本文提出三子空间解耦框架,通过分解多模态特征为公共、子模态共享和私人子空间,提升多模态情感分析的性能和鲁棒性。

Comments This study has been Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19542 2026-02-24 cs.CV

Vinedresser3D: Agentic Text-guided 3D Editing

Vinedresser3D: 基于代理的文本引导3D编辑

Yankuan Chi, Xiang Li, Zixuan Huang, James M. Rehg

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 Vinedresser3D通过多模态大语言模型和潜在空间编辑技术实现高质量文本引导的3D编辑,提升编辑精度与一致性。

Comments CVPR 2026, Project website:https://vinedresser3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19285 2026-02-24 cs.CV

MRI Contrast Enhancement Kinetics World Model

MRI对比增强动力学世界模型

Jindi Kong, Yuting He, Cong Xia, Rongjun Ge, Shuo Li

机构 * Case Western Reserve University(凯斯西储大学) Jiangsu Cancer Hospital(江苏癌症医院) Southeast University(东南大学)

AI总结 本文提出MRI CEKWorld模型,通过时空一致性学习解决MRI对比增强动力学建模中的时间连续性和内容一致性问题。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19180 2026-02-24 cs.CV

VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery

基于扩散模型的人体网格恢复中的群体偏好对齐

Wenhao Shen, Hao Wang, Wanqi Yin, Fayao Liu, Xulei Yang, Chao Liang, Zhongang Cai, Guosheng Lin

机构 * Nanyang Technological University(南洋理工大学) HKUST(GZ)(香港科技大学(广州)) SenseTime Research(商汤科技研究院) A*STAR(新加坡科技研究局)

AI总结 本文提出了一种基于群体偏好的扩散模型微调框架,通过双内存增强的批评代理生成质量评分,提升人体网格恢复的物理合理性和图像一致性。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19170 2026-02-24 cs.CV

BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality Assessment

BriMA:基于多模态适应的持续动作质量评估

Kanglei Zhou, Chang Li, Qingyi Pan, Liyuan Wang

机构 * Department of Psychological and Cognitive Sciences, 2 Department of Statistics and Data Science, Tsinghua University(1 心理学与认知科学系,2 统计学与数据科学系,清华大学)

AI总结 BriMA通过桥接模态适应方法,在模态缺失条件下提升多模态持续动作质量评估的性能。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19140 2026-02-24 cs.CV cs.LG

CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion

CaReFlow:循环适应修正流用于多模态融合

Sijie Mai, Shiqin Han

机构 * South China Normal University(华南师范大学)

AI总结 CaReFlow通过循环适应修正流实现多模态融合,解决模态间隙问题,提升分布对齐和特征转换的鲁棒性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19089 2026-02-24 cs.CV cs.GR cs.LG

Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic Sampling

Ani3DHuman:基于自引导随机采样的逼真3D人体动画

Qi Sun, Can Wang, Jiaxiang Shang, Yingchun Liu, Jing Liao

机构 * City University of Hong Kong(香港城市大学)

AI总结 Ani3DHuman通过结合运动学动画与视频扩散先验,利用自引导随机采样实现逼真3D人体动画生成。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19064 2026-02-24 cs.CV

L3DR: 3D-aware LiDAR Diffusion and Rectification

L3DR:基于3D感知的LiDAR扩散与校正

Quan Liu, Xiaoqin Zhang, Ling Shao, Shijian Lu

机构 * Nanyang Technological University(南洋理工大学) Zhejiang University of Technology(浙江工业大学) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 L3DR通过3D-aware的LiDAR扩散和校正框架,在3D空间中消除RV伪影并恢复局部几何结构,实现更真实的3D几何生成。

Comments In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19063 2026-02-24 cs.CV

Direction-aware 3D Large Multimodal Models

具有方向感知的3D大多模态模型

Quan Liu, Weihao Xuan, Junjue Wang, Naoto Yokoya, Ling Shao, Shijian Lu

AI总结 本文提出方向感知的3D大多模态模型,通过补充自身姿态并改进点云数据对齐,提升3D多模态模型的性能和通用性。

Comments In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06391 2026-02-24 cs.CV

Object-WIPER : Training-Free Object and Associated Effect Removal in Videos

Object-WIPER : 无需训练的对象及关联效果视频去除

Saksham Singh Kushwaha, Sayan Nag, Yapeng Tian, Kuldeep Kulkarni

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) Adobe Research(Adobe研究院)

AI总结 Object-WIPER通过预训练的文本到视频扩散模型实现无需训练的对象及关联效果视频去除,通过创新的去噪方法和新指标在DAVIS和WIPER-Bench上取得优越性能。

Comments Accepted to CVPR 2026. Project Page: https://sakshamsingh1.github.io/object_wiper_webpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17514 2026-02-24 cs.CV

Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection

基础模型先验增强特征空间中的目标聚焦以实现无源目标检测

Sairam VCR, Rishabh Lalla, Aveen Dayal, Tejal Kulkarni, Anuj Lalla, Vineeth N Balasubramanian, Muhammad Haris Khan

机构 * IIT Hyderabad(印度海得拉尔理工学院) MBZUAI, Abu Dhabi(阿布扎赫尔MBZUAI) UC San Diego(圣地亚哥大学) IIT Jodhpur(乔普尔理工学院) Microsoft Research India(微软印度研究院)

AI总结 FALCON-SFOD通过增强特征空间中的目标聚焦,利用基础模型先验和噪声鲁棒伪标签方法,提升无源目标检测在领域偏移下的性能。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15690 2026-02-24 cs.CV cs.CL

MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping

通过动态专家跳过加速混合专家多模态大语言模型

Yushi Huang, Zining Wang, Zhihang Yuan, Yifu Ding, Ruihao Gong, Jinyang Guo, Xianglong Liu, Jun Zhang

机构 * Hong Kong University of Science and Technology(香港科技大学) Beihang University(北航) Peking University(北京大学)

AI总结 MoDES通过动态专家跳过机制提升多模态大语言模型的推理效率和准确性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08318 2026-02-24 cs.CV

LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation

LinVideo: 一种面向高效视频生成的O(n)注意力后训练框架

Yushi Huang, Xingtong Ge, Ruihao Gong, Chengtao Lv, Jun Zhang

机构 * Hong Kong University of Science and Technology(香港科技大学) Beihang University(北航) Sensetime Research(商汤科技研究院) Nanyang Technological University(南洋理工大学)

AI总结 LinVideo提出一种高效的无数据后训练框架,通过选择性迁移将部分自注意力模块替换为线性注意力,从而在保持性能的同时显著提升视频生成效率。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10065 2026-02-24 cs.CV

MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second

MoVieS: 一秒钟内基于运动感知的4D动态视角合成

Chenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu, Tao Hu, Honglei Yan, Katerina Fragkiadaki, Yadong Mu

机构 * Peking University(北京大学) ByteDance(字节跳动) Carnegie Mellon University(卡内基梅隆大学)

AI总结 MoVieS通过一秒钟内从单目视频重建4D动态场景,实现外观、几何和运动的统一建模,并支持多种零样本应用。

Comments Project page: https://chenguolin.github.io/projects/MoVieS; Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03584 2026-02-24 cs.CV cs.AI

RDFC-GAN: RGB-Depth Fusion CycleGAN for Indoor Depth Completion

RDFC-GAN:基于RGB-深度融合的循环GAN用于室内深度补全

Haowen Wang, Zhengping Che, Yufan Yang, Mingyuan Wang, Zhiyuan Xu, Xiuquan Qiao, Mengshi Qi, Feifei Feng, Jian Tang

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China(网络与交换技术国家重点实验室,北京邮电大学,中国) Midea Group, China(美的集团,中国) School of Computer Science, Beijing University of Posts and Telecommunications, China(计算机科学学院,北京邮电大学,中国)

AI总结 RDFC-GAN通过融合RGB和深度图像,利用循环GAN和自适应融合模块提升室内深度补全效果。

Comments Haowen Wang and Zhengping Che are with equal contributions. Paper accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). An earlier version has been accepted by CVPR 2022 (arXiv:2203.10856). arXiv admin note: text overlap with arXiv:2203.10856

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence (Volume: 46, Issue: 11, November 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19024 2026-02-24 cs.CV

Towards Calibrating Prompt Tuning of Vision-Language Models

面向视觉-语言模型提示微调的校准

Ashshak Sharifdeen, Fahad Shamshad, Muhammad Akhtar Munir, Abhishek Basu, Mohamed Insaf Ismithdeen, Jeyapriyan Jeyamohan, Chathurika Sewwandi Silva, Karthik Nandakumar, Muhammad Haris Khan

机构 * Mohamed bin Zayed University of AI(Mohamed bin Zayed人工智能大学) University of Colombo(科伦坡大学) Michigan State University(密歇根州立大学)

AI总结 本文提出一种校准框架,通过引入均值-方差边际惩罚和文本矩匹配损失,提升视觉-语言模型提示微调的预测可靠性与置信度校准。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18977 2026-02-24 cs.CV

Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding

Frame2Freq: 用于细粒度视频理解的频谱适配器

Thinesh Thiyakesan Ponbagavathi, Constantin Seibold, Alina Roitberg

机构 * Institute for Artificial Intelligence, University of Stuttgart(斯图加特大学人工智能研究所) University Hospital Heidelberg, Diagnostic and Interventional Radiology(海德堡大学医院放射诊断与介入科) Intelligent Assistive Systems Lab, University of Hildesheim(希尔德斯海姆大学智能辅助系统实验室)

AI总结 Frame2Freq通过频谱编码提升细粒度视频理解,利用FFT和频率带嵌入改进动作识别性能。

Comments Accepted to CVPR 2026 (Main Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18863 2026-02-24 eess.IV cs.CV cs.LG cs.MM

TIACam: Text-Anchored Invariant Feature Learning with Auto-Augmentation for Camera-Robust Zero-Watermarking

TIACam: 基于自增强的文本锚定不变特征学习用于抗相机零水印

Abdullah All Tanvir, Agnibh Dasgupta, Xin Zhong

机构 * Department of Computer Science University of Nebraska Omaha(计算机科学系 内布拉斯加大学奥马哈分校)

AI总结 TIACam通过文本锚定不变特征学习和自增强技术,实现抗相机重拍的零水印系统,提升特征稳定性和水印提取准确性。

Comments This paper is accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18842 2026-02-24 cs.CV

Detecting AI-Generated Forgeries via Iterative Manifold Deviation Amplification

通过迭代流形偏差放大检测AI生成的伪造

Jiangling Zhang, Shuxuan Gao, Bofan Liu, Siqiang Feng, Jirui Huang, Yaxiong Chen, Ziyu Chen

机构 * Wuhan University of Technology(武汉理工大学)

AI总结 提出IFA-Net通过迭代流形偏差放大技术,实现对AI生成伪造的高精度检测,提升篡改区域定位的准确性和泛化能力。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18811 2026-02-24 cs.CV

Learning Multi-Modal Prototypes for Cross-Domain Few-Shot Object Detection

跨域少样本目标检测中的多模态原型学习

Wanqi Wang, Jingcai Guo, Yuxiang Cai, Zhi Chen

机构 * University of Chinese Academy of Sciences(中国科学院大学) The Hong Kong Polytechnic University(香港理工大学) Zhejiang University(浙江大学) The University of Southern Queensland(昆士兰大学)

AI总结 本文提出LMP方法,通过结合文本和视觉信息,提升跨域少样本目标检测的精度和性能。

Comments Accepted to CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16412 2026-02-24 cs.CV

ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding

ReMoRa:基于精细运动表示的多模态大语言模型用于长视频理解

Daichi Yashima, Shuhei Kurita, Yusuke Oda, Komei Sugiura

机构 * Keio University(庆应大学) NII(日本信息处理学会) NII LLMC(日本信息处理学会语言模型中心)

AI总结 ReMoRa通过精细运动表示实现长视频理解,有效压缩视频数据并提升多模态大语言模型性能。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09217 2026-02-24 cs.RO cs.CV stat.AP

Perception Characteristics Distance: Measuring Stability and Robustness of Perception System in Dynamic Conditions under a Certain Decision Rule

感知特性距离:在特定决策规则下动态条件下感知系统稳定性与鲁棒性测量

Boyu Jiang, Liang Shi, Zhengzhi Lin, Lanxin Xiang, Loren Stowe, Feng Guo

机构 * Department of Statistics, Virginia Tech(弗吉尼亚理工学院统计学系) Virginia Tech Transportation Institute(弗吉尼亚理工学院交通研究所)

AI总结 提出感知特性距离(PCD)作为衡量动态条件下感知系统稳定性与鲁棒性的新指标,并通过SensorRainFall数据集验证其有效性。

Comments This paper has been accepted to the CVPR 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00578 2026-02-24 cs.CV cs.GR

Speedy-Splat: Fast 3D Gaussian Splatting with Sparse Pixels and Sparse Primitives

Speedy-Splat: 通过稀疏像素和稀疏原语实现快速3D高斯点播

Alex Hanson, Allen Tu, Geng Lin, Vasu Singla, Matthias Zwicker, Tom Goldstein

机构 * University of Maryland, College Park(马里兰大学 College Park分校)

AI总结 Speedy-Splat通过优化渲染管线和引入剪枝技术,显著提升了3D高斯点播的渲染速度,模型大小和训练时间也得到优化。

Comments CVPR 2025, Project Page: https://speedysplat.github.io/

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 21537-21546

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18322 2026-02-23 cs.CV

Unifying Color and Lightness Correction with View-Adaptive Curve Adjustment for Robust 3D Novel View Synthesis

统一颜色和亮度校正与视图自适应曲线调整以实现鲁棒的3D新视角合成

Ziteng Cui, Shuhong Liu, Xiaoyu Dong, Xuangeng Chu, Lin Gu, Ming-Hsuan Yang, Tatsuya Harada

机构 * University of Tokyo(东京大学) Tohoku University(东北大学) University of California at Merced(加州默塞德大学) Google DeepMind(谷歌DeepMind) RIKEN AIP(理化学研究所AIP)

AI总结 Luminance-GS++通过结合全局视图自适应亮度调整和局部像素级残差细化,实现鲁棒的3D新视角合成,提升重建保真度并保持实时渲染效率。

Comments Journal extension version of CVPR 2025 paper: arXiv:2504.01503

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19026 2026-02-20 cs.CV

MeGA: Hybrid Mesh-Gaussian Head Avatar for High-Fidelity Rendering and Head Editing

MeGA: 混合网格-高斯头像用于高保真渲染和头像编辑

Cong Wang, Di Kang, He-Yi Sun, Shen-Han Qian, Zi-Xuan Wang, Linchao Bao, Song-Hai Zhang

机构 * Tsinghua University(清华大学) Tencent(腾讯) Technical University of Munich(慕尼黑技术大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 MeGA通过混合网格和高斯表示,实现高保真头像渲染与编辑,提升面部和头发的细节与真实感。

Comments Accepted by CVPR 2025. Project page: https://conallwang.github.io/MeGA_Pages/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16537 2026-02-19 cs.CV cs.AI cs.CL cs.RO

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

RoboSpatial: 教授机器人2D和3D视觉-语言模型空间理解能力

Chan Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree, Yu Su, Stan Birchfield

机构 * The Ohio State University(俄亥俄州立大学) NVIDIA(英伟达)

AI总结 RoboSpatial通过构建大规模空间理解数据集,提升机器人2D和3D视觉-语言模型的空间推理能力。

Comments CVPR 2025 (Oral); Project Website: https://chanh.ee/RoboSpatial

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12385 2026-02-17 cs.CV

Car-1000: A New Large Scale Fine-Grained Visual Categorization Dataset

Car-1000:一个新的大规模细粒度视觉分类数据集

Yutao Hu, Sen Li, Jincheng Yan, Wenqi Shao, Xiaoyan Luo

机构 * Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education(新一代人工智能技术及交叉应用国家重点实验室(东南大学)) School of Astronautics, Beihang University(北航航天学院) Shanghai AI Laboratory(上海人工智能实验室)

AI总结 Car-1000是一个大规模细粒度视觉分类数据集,包含1000种不同汽车型号,旨在提升自动驾驶和交通监控等领域的识别性能。

Comments accepted to The Eleventh Workshop on Fine-Grained Visual Categorization in CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15605 2026-02-17 cs.CV cs.LG

Efficiently Assemble Normalization Layers and Regularization for Federated Domain Generalization

高效组装规范化层和正则化以实现联邦域泛化

Khiem Le, Long Ho, Cuong Do, Danh Le-Phuoc, Kok-Seng Wong

机构 * Department of Computer Science and Engineering, University of Notre Dame(Notre Dame 大学计算机科学与工程系) College of Engineering and Computer Science, VinUniversity(Vin大学工程与计算机科学学院) Open Distributed Systems, Technical University Berlin(柏林技术大学开放分布式系统)

AI总结 gPerXAN通过个性化规范化和引导正则化实现联邦域泛化的高效组装,提升模型在域偏移下的泛化能力。

Comments CVPR'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12393 2026-02-16 cs.CV cs.AI cs.LG

Reproducing DragDiffusion: Interactive Point-Based Editing with Diffusion Models

重现DragDiffusion:基于扩散模型的交互式点编辑

Ali Subhan, Ashir Raza

机构 * Faculty of Computer and Information Science(计算机与信息科学学院) University of Ljubljana(卢布尔雅那大学)

AI总结 本文重现了DragDiffusion方法,验证了其在不同超参数下的可重复性,并发现其性能对优化时间步和特征层敏感。

Comments 16 pages, 8 figures. Reproducibility study of DragDiffusion (CVPR 2024). Submitted to TMLR Reproducibility Challenge. Code available on GitHub

详情

展开后加载摘要…

URL PDF HTML 收藏