arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11875
2603.20808 2026-03-24 cs.CV cs.LG

Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models

预测正则化对抗多模态大语言模型中的视觉表征退化

Enguang Wang, Qiang Wang, Yuanchen Wu, Ke Yan, Xinbin Yuan, Shouhong Ding, Xialei Liu, Ming-Ming Cheng

机构 * NKIARI VCIP, CS, Nankai University(VCIP计算机科学系,南开大学) AAIS, Nankai University(AAIS,南开大学) Tencent Youtu Lab(腾讯优设实验室)

AI总结 本文研究多模态大语言模型中的视觉表征退化问题,提出预测正则化方法以维持视觉表征,提升视觉语言性能。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20782 2026-03-24 cs.CV

MEMO: Human-like Crisp Edge Detection Using Masked Edge Prediction

MEMO: 通过掩码边缘预测实现类人清晰边缘检测

Jiaxin Cheng, Yue Wu, Yicong Zhou

机构 * Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系)

AI总结 本文提出MEMO模型,通过设计训练和推理策略实现清晰边缘检测,利用交叉熵损失在合成数据集上预训练,再通过轻量模块微调,最终通过逐步预测策略生成更精确的边缘。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20755 2026-03-24 cs.CV cs.AI

Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping

通过动态补丁采样和块跳过实现高效的扩散变换器微调

Sunghyun Park, Jeongho Kim, Hyoungwoo Park, Debasmit Das, Sungrack Yun, Munawar Hayat, Jaegul Choo, Fatih Porikli, Seokeon Choi

机构 * Qualcomm AI Research(高通人工智能研究) KAIST(韩国科学技术院)

AI总结 本文提出DiT-BlockSkip框架,通过动态补丁采样和块跳过减少内存使用,提升扩散变换器微调效率,实现设备端部署。

Comments Accepted to CVPR 2026; 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20741 2026-03-24 cs.CV

CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-Calibration

CTCal: 通过跨时间步自校准重新思考文本到图像扩散模型

Xiefan Guo, Xinzhu Ma, Haiyu Zhang, Di Huang

机构 * State Key Laboratory of Complex and Critical Software Environment(复杂与关键软件环境国家重点实验室) School of Computer Science and Engineering(计算机科学与工程学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

AI总结 本文提出CTCal,通过利用早期时间步的准确文本-图像对齐来校准后期时间步的表示学习,提升文本到图像生成的对齐精度。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20739 2026-03-24 cs.CV

Mamba Learns in Context: Structure-Aware Domain Generalization for Multi-Task Point Cloud Understanding

Mamba在上下文中学习:面向多任务点云理解的结构感知领域泛化

Jincen Jiang, Qianyu Zhou, Yuhang Li, Kui Su, Meili Wang, Jian Chang, Jian Jun Zhang, Xuequan Lu

机构 * Bournemouth University(伯恩茅斯大学) Jilin University(吉林大学) The University of Western Australia(西澳大学) Hangzhou City University(杭州城市大学) Northwest A&F University(西北农林科技大学)

AI总结 本文提出SADG框架,通过结构感知序列化和层次领域感知建模提升多任务点云领域的结构一致性与性能。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20721 2026-03-24 cs.CV

Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark

跨模态模糊对齐网络用于文本-空中人检索及一个大规模基准

Yifei Deng, Chenglong Li, Yuyang Zhang, Guyue Hu, Jin Tang

机构 * State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology(光电信息采集与防护技术国家重点实验室) School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) School of Artificial Intelligence, Anhui University(安徽大学人工智能学院) The University of Hong Kong(香港大学)

AI总结 本文提出跨模态模糊对齐网络,通过模糊逻辑量化token级可靠性,结合地面图像作为桥梁代理,提升文本与空中图像的语义对齐鲁棒性,并构建了大规模基准数据集AERI-PEDES。

Comments Accepted by CVPR 2026 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20708 2026-03-24 cs.CV

High-Quality and Efficient Turbulence Mitigation with Events

高质高效湍流抑制方法

Xiaoran Zhang, Jian Ding, Yuxing Duan, Haoyue Liu, Gang Chen, Yi Chang, Luxin Yan

机构 * State Key Laboratory of Multispectral Information Intelligent Processing Technology(多谱信息智能处理技术国家重点实验室) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

AI总结 本文提出EHETM方法,利用事件相机的特性,通过事件极性变化和事件管约束实现高效湍流抑制,提升恢复质量并减少数据开销和系统延迟。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20611 2026-03-24 cs.CV

GaussianPile: A Unified Sparse Gaussian Splatting Framework for Slice-based Volumetric Reconstruction

GaussianPile:一种统一的稀疏高斯散射框架用于基于切片的体积分离重建

Di Kong, Yikai Wang, Wenjie Guo, Yifan Bu, Boya Zhang, Yuexin Duan, Xiawei Yue, Wenbiao Du, Yiman Zhong, Yuwen Chen, Cheng Ma

机构 * Tsinghua University(清华大学) Zhongguancun Academy(中关村学院) Beijing Normal University(北京师范大学) Nankai University(南开大学) Beijing Institute of Technology(北京理工大学) Beihang University(北航)

AI总结 本文提出GaussianPile框架,结合3D高斯散射与成像系统感知的聚焦模型,通过切片感知堆叠策略、可微投影算子和紧凑编码优化流程,实现高效体积分离重建,提升压缩效率与诊断精度。

Comments Accepted by IEEE/CVF Conference on Computer Vision and Pattern Recognition 2026 (CVPR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17655 2026-03-24 cs.CV cs.AI

Interpretable Cross-Domain Few-Shot Learning with Rectified Target-Domain Local Alignment

可解释的跨领域少样本学习与修正的目标域局部对齐

Yaze Zhao, Yixiong Zou, Yuhua Li, Ruixuan Li

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)

AI总结 本文提出CC-CDFSL方法,通过循环一致性解决CLIP-based CDFSL中的局部对齐问题,提升局部视觉语言对齐和可解释性,实现SOTA性能。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12918 2026-03-24 cs.CV

VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose Estimation

VIRD:通过双轴变换实现视图不变表示的跨视图姿态估计

Juhye Park, Wooju Lee, Dasol Hong, Changki Sung, Youngwoo Seo, Dongwan Kang, Hyun Myung

机构 * Urban Robotics Lab, School of Electrical Engineering, KAIST(首尔国立大学电气工程学院智能城市机器人实验室) Hanwha Aerospace(韩华航空航天)

AI总结 本文提出VIRD方法,通过双轴变换构建视图不变表示,解决地面与卫星视图间显著视角差异问题,提升跨视图姿态估计精度。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08536 2026-03-24 cs.CV

SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attribution

SWIFT: 基于滑动窗口的少样本训练自由生成视频属性识别

Chao Wang, Zijin Yang, Yaofei Wang, Yuang Qi, Weiming Zhang, Nenghai Yu, Kejiang Chen

机构 * University of Science and Technology of China(中国科学技术大学) Hefei University of Technology(合肥工业大学)

AI总结 本文提出SWIFT方法,通过滑动窗口实现视频属性识别,无需额外训练即可在多种生成模型上达到90%以上的准确率,支持零样本识别。

Comments 8 pages. Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03744 2026-03-24 cs.CV

DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation

DAGE:用于高效和细粒度几何估计的双流架构

Tuan Duc Ngo, Jiahui Huang, Seoung Wug Oh, Kevin Blackburn-Matzen, Evangelos Kalogerakis, Chuang Gan, Joon-Young Lee

机构 * UMass Amherst(马萨诸塞大学阿姆赫斯特分校) Adobe Research(Adobe研究) TU Crete(希腊技术大学)

AI总结 DAGE提出双流变压器架构,通过分离全局一致性与细节,实现高效且细粒度的几何估计与相机姿态估计,支持高分辨率和长序列输入。

Comments CVPR 2026. Project page: https://ngoductuanlhp.github.io/dage-site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00431 2026-03-24 cs.CV cs.AI

Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models

面向层次视觉识别的分类意识表示对齐

Hulingxiao He, Zhi Tan, Yuxin Peng

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机系王轩研究所)

AI总结 本文提出TARA方法,通过生物基础模型的层次对比学习将分类知识注入大模态模型,提升层次视觉识别中对已知和新类别的识别性能。

Comments Published as a conference paper at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21499 2026-03-24 cs.CV

Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel Flow

Easy3E: 通过校正体素流实现的前馈3D资产编辑

Shimin Hu, Yuanyi Wei, Fei Zha, Yudong Guo, Juyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出基于TRELLIS生成框架的高效前馈3D编辑方法,通过单视角编辑实现3D模型修改,解决训练自由2D编辑适应结构化3D表示及压缩3D特征外观保真度瓶颈问题。

Comments CVPR 2026, Project Page: https://ustc3dv.github.io/Easy3E/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10744 2026-03-24 cs.AI cs.CV

Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration

探索与长期记忆:一个基准和基于多模态LLM的强化学习框架用于具身探索

Sen Wang, Bangwei Liu, Zhenkun Gao, Lizhuang Ma, Xuhong Wang, Yuan Xie, Xin Tan

机构 * East China Normal University(东华大学) Shanghai AI Laboratory(上海人工智能实验室)

AI总结 本文提出LMEE框架,通过多模态LLM强化学习促进终身学习,构建LMEE-Bench基准评估具身探索过程与结果,采用MemoryExplorer方法提升记忆检索与主动探索能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05175 2026-03-24 cs.CV

VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice

VideoAuto-R1:通过一次推理、两次回答实现视频自动推理

Shuming Liu, Mingchen Zhuge, Changsheng Zhao, Jun Chen, Lemeng Wu, Zechun Liu, Chenchen Zhu, Zhipeng Cai, Chong Zhou, Haozhe Liu, Ernie Chang, Saksham Suri, Hongyu Xu, Qi Qian, Wei Wen, Balakrishnan Varadarajan, Zhuang Liu, Hu Xu, Florian Bordes, Raghuraman Krishnamoorthi, Bernard Ghanem, Vikas Chandra, Yunyang Xiong

机构 * Meta AI King Abdullah University of Science and Technology (KAUST)(卡布斯大学) Princeton University(普林斯顿大学)

AI总结 本文提出VideoAuto-R1框架,通过一次推理两次回答策略提升视频理解效率,实现准确率和效率的双重提升。

Comments Accepted to CVPR 2026. Project page: https://ivul-kaust.github.io/projects/videoauto-r1/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21778 2026-03-24 cs.CV

Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models

Scene-VLM:通过视觉-语言模型进行多模态视频场景分割

Nimrod Berman, Adam Botach, Emanuel Ben-Baruch, Shunit Haviv Hakimi, Asaf Gendler, Ilan Naiman, Erez Yosef, Igor Kviatkovsky

机构 * Ben-Gurion University(本·古里安大学) Amazon Prime Video(亚马逊Prime视频) Tel-Aviv University(特拉维夫大学)

AI总结 本文提出Scene-VLM,首个基于视觉-语言模型的视频场景分割框架,通过融合视觉与文本信息实现多模态推理,提升场景分割的准确性和可解释性。

Comments Accepted for publication at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19402 2026-03-24 cs.RO cs.CV cs.GR

Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface

Real2Edit2Real:通过3D控制界面生成机器人演示

Yujie Zhao, Hongwei Fan, Di Chen, Shengcong Chen, Liliang Chen, Xiaoqi Li, Guanghui Ren, Hao Dong

机构 * CFCS, School of Computer Science, Peking University(计算机学院,北京大学) PKU-AgiBot Lab(北大AgiBot实验室) AgiBot

AI总结 本文提出Real2Edit2Real框架,通过3D可编辑性与2D视觉数据结合,减少重复数据收集,提升空间泛化能力,实验表明其数据效率提升显著。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05959 2026-03-24 cs.CL cs.AI cs.CV

M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG

M4-RAG:大规模多语言多文化多模态检索增强生成

David Anugraha, Patrick Amadeus Irawan, Anshul Singh, En-Shiun Annie Lee, Genta Indra Winata

机构 * Stanford University(斯坦福大学) MBZUAI Indian Institute of Science(印度科学研究院) Ontario Tech University(安大略技术大学) University of Toronto(多伦多大学) Capital One

AI总结 M4-RAG提出一个覆盖42种语言、56种方言和189个国家的多模态大规模基准,通过构建8万多个文化多样化的图像-问题对,评估跨语言和模态的检索增强视觉问答性能,揭示模型大小与检索效果的不匹配问题。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03794 2026-03-24 cs.CV cs.AI cs.CL cs.LG

AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition

AdaptVision: 通过自适应视觉获取实现高效的视觉-语言模型

Zichuan Lin, Yicheng Liu, Yang Yang, Lvfang Tao, Deheng Ye

机构 * Tencent Hunyuan(腾讯文言)

AI总结 AdaptVision通过自适应视觉获取机制,结合强化学习与工具学习,提升视觉-语言模型在视觉问答任务中的效率与准确性。

Comments Accepted by CVPR 2026. Code and models are available at https://github.com/AdaptVision/AdaptVision

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00422 2026-03-24 cs.CV

PhysGen: Physically Grounded 3D Shape Generation for Industrial Design

PhysGen:面向工业设计的物理引导的3D形状生成

Yingxuan You, Chen Zhao, Hantao Zhang, Ming Xu, Pascal Fua

机构 * CVLab, EPFL(EPFL计算机视觉实验室)

AI总结 PhysGen提出一种结合物理引导的3D形状生成方法,通过物理感知正则化项和变分自编码器提升形状真实感,实验表明其优于单纯视觉合理性。

Comments Accepted to CVPR 2026. 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16407 2026-03-24 cs.RO

LAOF: Robust Latent Action Learning with Optical Flow Constraints

LAOF:基于光流约束的鲁棒潜在动作学习

Xizhou Bu, Jiexi Lyu, Fulei Sun, Ruichen Yang, Zhiqiang Ma, Wei Li

机构 * Fudan University(复旦大学) Northwestern Polytechnical University(西北工业大学)

AI总结 本文提出LAOF方法,通过利用光流作为动作驱动信号,学习鲁棒的潜在动作表示,以应对动作无关的干扰。实验表明,LAOF在下游模仿学习和强化学习任务中表现优异,尤其在标注稀缺条件下效果显著。

Comments CVPR 2026; Project page: https://github.com/XizoB/LAOF

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15130 2026-03-24 cs.GR cs.AI cs.CV

Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control

通过零样本相机控制驯服视频模型以实现3D和4D生成

Chenxi Song, Yanming Yang, Tong Zhao, Ruibo Li, Chi Zhang

机构 * AGI Lab, Westlake University(西溪大学AGI实验室) Nanyang Technological University(南洋理工大学)

AI总结 本文提出WorldForge框架,通过推理时的三重组件实现零样本3D/4D生成,解决视频扩散模型在空间任务中的控制不足问题,提升视觉真实性与轨迹一致性。

Comments Accepted to CVPR 2026. Project Webpage: https://worldforge-agi.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13677 2026-03-24 cs.CV cs.AI cs.LG cs.MM

HeCoFuse: Cross-Modal Complementary V2X Cooperative Perception with Heterogeneous Sensors

HeCoFuse: 跨模态互补V2X协作感知与异构传感器

Chuheng Wei, Ziye Qin, Walter Zimmer, Guoyuan Wu, Matthew J. Barth

机构 * College of Engineering, Center for Environmental Research and Technology, University of California at Riverside(加州大学河滨分校工程学院、环境研究与技术中心) School of Transportation and Logistics, Southwest Jiaotong University(西南交通大学交通运输与物流学院) Chair of Robotics, Artificial Intelligence and Real-time Systems, TUM School of Computation, Information and Technology, Technical University of Munich(慕尼黑技术大学计算机、信息与技术学院机器人、人工智能与实时系统教授职位)

AI总结 HeCoFuse提出一种统一框架,通过通道和空间注意力机制实现异构传感器下的跨模态特征融合,提升协作感知的可靠性与性能,实验显示其在TUMTraf-V2X数据集上达到43.22%的3D mAP,优于基线方法。

Comments Ranked first in CVPR DriveX workshop TUM-Traf V2X challenge. Accepted by ITSC2025

Journal ref Proceedings of the 2025 IEEE 28th International Conference on Intelligent Transportation Systems (ITSC), pp. 1214-1221, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23495 2026-03-24 cs.CV

Embedding Shift Dissection on CLIP: Effects of Augmentations on VLM's Representation Learning

在CLIP上嵌入位移分析:不同增强对VLM表示学习的影响

Ashim Dahal, Saydul Akbar Murad, Nick Rahimi

机构 * University of Southern Mississippi(密西西比大学)

AI总结 研究分析不同增强技术对CLIP嵌入位移的影响,探讨增强对视觉语言模型表示学习的机械可解释性影响。

Comments accepted at MIV at CVPR 2025

Journal ref 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20470 2026-03-24 cs.AI

DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image Generation

DiffGraph: 一种自动化代理驱动的模型融合框架用于真实场景的文本到图像生成

Zhuoling Li, Hossein Rahmani, Jiarui Zhang, Yu Xue, Majid Mirmehdi, Jason Kuen, Jiuxiang Gu, Jun Liu

机构 * Lancaster University(兰卡斯特大学) University of Bristol(布里斯托大学) Adobe Research(Adobe研究院)

AI总结 DiffGraph通过自动化整合在线专家资源,灵活融合不同模型以满足多样化的真实用户需求,提升了文本到图像生成的效能。

Comments CVPR

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20448 2026-03-24 cs.CV eess.IV

Thermal is Always Wild: Characterizing and Addressing Challenges in Thermal-Only Novel View Synthesis

热总是野性:在仅热成像的新型视角合成中的特征刻画与挑战应对

M. Kerem Aydin, Vishwanath Saragadam, Emma Alexander

机构 * Northwestern University(西北大学) University of California, Riverside(加州大学河滨分校)

AI总结 本文针对热成像在新型视角合成中的难点,提出轻量预处理与溅射流程,提升动态范围并稳定光照度,实现先进性能。

Comments To be published at CVPR, 2026. 15 Pages, 29 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20403 2026-03-24 cs.CV

FAAR: Efficient Frequency-Aware Multi-Task Fine-Tuning via Automatic Rank Selection

FAAR:通过自动排名选择实现高效的频率感知多任务微调

Maxime Fontana, Michael Spratling, Miaojing Shi

机构 * King’s College London(伦敦国王学院) University of Luxembourg(卢森堡大学) Tongji University(同济大学)

AI总结 本文提出FAAR方法,通过性能驱动的排名缩减和任务频谱金字塔解码器,在多任务学习中提升准确性和效率,相比传统方法减少参数达9倍。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20323 2026-03-24 cs.CV

NCSTR: Node-Centric Decoupled Spatio-Temporal Reasoning for Video-based Human Pose Estimation

NCSTR:基于节点的解耦时空推理用于视频人体姿态估计

Quang Dang Huynh, Xuefei Yin, Andrew Busch, Hugo G. Espinosa, Alan Wee-Chung Liew, Matthew T. O. Worsey, Yanming Zhu

机构 * Griffith University(格里菲斯大学)

AI总结 本文提出NCSTR框架,通过融合视觉、时间和结构推理,解决视频人体姿态估计中的模糊、遮挡和复杂时空动态问题,通过双分支解耦时空注意力图和节点空间专家融合模块提升姿态估计精度。

Journal ref CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17487 2026-03-24 cs.CV

Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models

智能下放:探索小型多模态模型中感知与推理的瓶颈

Mark Endo, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

AI总结 本文研究小型多模态模型中感知与推理能力的瓶颈,通过分析LLM容量下降对多模态能力的影响,提出视觉提取调优方法,提升效率与性能。

Comments CVPR 2026, website at https://web.stanford.edu/~markendo/projects/downscaling_intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏