arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

International Journal of Computer Vision · 期刊 · Computer Vision

至 收录 556
2411.02057 2026-01-27 cs.CV

Exploiting Unlabeled Data with Multiple Expert Teachers for Open Vocabulary Aerial Object Detection and Its Orientation Adaptation

利用多专家教师挖掘未标注数据用于开放词汇空中物体检测及其方向适应

Yan Li, Weiwei Guo, Xue Yang, Ning Liao, Shaofeng Zhang, Yi Yu, Wenxian Yu, Junchi Yan

机构 * Tongji University(同济大学) Southeast University(东南大学)

AI总结 本文提出CastDet框架,通过多专家教师和动态标签队列实现开放词汇空中物体检测及其方向适应,提升新型物体识别与分类能力。

Comments Accepted by International Journal of Computer Vision (IJCV'26)

URL PDF HTML 收藏
2601.13148 2026-01-21 cs.CV cs.HC

ICo3D: An Interactive Conversational 3D Virtual Human

ICo3D: 一种交互式对话式3D虚拟人

Richard Shaw, Youngkyoon Jang, Athanasios Papaioannou, Arthur Moreau, Helisa Dhamo, Zhensong Zhang, Eduardo Pérez-Pellitero

机构 * Noah’s Ark Lab(诺亚 Ark 实验室)

AI总结 ICo3D通过动态高斯模型和LLM实现逼真3D虚拟人,支持实时对话与交互。

Comments Accepted by International Journal on Computer Vision (IJCV). Project page: https://ico3d.github.io/. This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this article is published in International Journal of Computer Vision and is available online at https://doi.org/10.1007/s11263-025-02725-8

URL PDF HTML 收藏
2406.12805 2026-01-16 cs.CV

AITTI: Learning Adaptive Inclusive Token for Text-to-Image Generation

AITTI: 学习自适应包容性令牌以生成文本到图像

Xinyu Hou, Xiaoming Li, Chen Change Loy

机构 * Nanyang Technological University(南洋理工大学)

AI总结 AITTI通过学习自适应包容性令牌,有效缓解文本到图像生成中的刻板印象偏见,无需显式属性指定或先验知识。

Comments Accepted by IJCV

URL PDF HTML 收藏
2601.07117 2026-01-13 cs.CV cs.AI

Few-shot Class-Incremental Learning via Generative Co-Memory Regularization

基于生成共记忆正则化的少样本类增量学习

Kexin Bao, Yong Li, Dan Zeng, Shiming Ge

机构 * Institute of Information Engineering(信息工程研究所) Chinese Academy of Sciences(中国科学院) School of Cyber Security(网络安全学院) University of Chinese Academy of Sciences(中国科学院大学) Department of Communication Engineering(通信工程系)

AI总结 本文提出基于生成共记忆正则化的少样本类增量学习方法,通过微调生成编码器和构建类记忆来提升模型在少量样本下的识别准确率,同时减少灾难性遗忘和过拟合。

Comments Accepted by International Journal on Computer Vision (IJCV)

URL PDF HTML 收藏
2506.04115 2026-01-13 cs.CV

Multi-view Surface Reconstruction Using Normal and Reflectance Cues

多视角表面重建使用法线和反照率线索

Robin Bruneau, Baptiste Brument, Yvain Quéau, Jean Mélou, François Bernard Lauze, Jean-Denis Durou, Lilian Calvet

机构 * University of Zurich(苏黎世大学) Université de Toulouse(图卢兹大学) CNRS, UNICAEN, ENSICAEN, Normandie Université(法国国家科学研究中心、UNICAEN、ENSICAEN、诺曼底大学) FittingBox(FittingBox公司) University of Copenhagen(哥本哈根大学)

AI总结 本文提出了一种结合多视角法线和反照率地图的表面重建框架,通过参数化辐射向量实现高保真3D重建,尤其在复杂材质和可见性挑战下表现优异。

Comments 22 pages, 15 figures, 11 tables. Accepted to IJCV. A thorough qualitative and quantitive study is available in the supplementary material at https://drive.google.com/file/d/1KDfCKediXNP5Os954TL_QldaUWS0nKcD/view?usp=drive_link. The project page can be accessed via https://robinbruneau.github.io/publications/rnb_neus2.html. The source code is available at https://github.com/RobinBruneau/RNb-NeuS2

URL PDF HTML 收藏
2601.05244 2026-01-09 cs.CV

GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation

GREx: 通用指称表达分割、理解和生成

Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, Yu-Gang Jiang

机构 * Fudan University(复旦大学)

AI总结 GREx提出通用指称表达分割、理解和生成框架,扩展至多目标和无目标表达,引入gRefCOCO数据集和ReLA方法,提升复杂关系建模能力。

Comments IJCV, Project Page: https://henghuiding.com/GREx/

URL PDF HTML 收藏
2410.11041 2026-01-09 cs.CV

Beyond Fixed Topologies: Unregistered Training and Comprehensive Evaluation Metrics for 3D Talking Heads

超越固定拓扑:用于3D说话人脸的无注册训练和综合评估指标

Federico Nocentini, Thomas Besnier, Claudio Ferrari, Sylvain Arguillere, Mohamed Daoudi, Stefano Berretti

机构 * MICC, University of Florence(佛罗伦萨大学MICC) CRIStAL, University of Lille, CNRS, Centrale Lille(里尔大学CRIStAL) IMT Nord Europe, University of Lille(里尔大学IMT Nord Europe) Dept. of Information Engineering and Mathematics, University of Siena(锡耶纳大学信息工程与数学系) Laboratoire Paul Painlevé, University of Lille, CNRS(里尔大学Paul Painlevé实验室)

AI总结 本文提出了一种能够处理任意拓扑的3D说话人脸生成方法,通过热扩散预测鲁棒特征,并引入新的评估指标以提升唇同步效果。

Comments https://fedenoce.github.io/scantalk/

Journal ref International Journal of Computer Vision 2025

URL PDF HTML 收藏
2412.18342 2026-01-08 cs.CV cs.LG eess.IV

Mitigating Label Noise using Prompt-Based Hyperbolic Meta-Learning in Open-Set Domain Generalization

通过基于提示的双曲元学习缓解标签噪声以实现开放集领域泛化

Kunyu Peng, Di Wen, M. Saquib Sarfraz, Yufan Chen, Junwei Zheng, David Schneider, Kailun Yang, Jiamin Wu, Alina Roitberg, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Hunan University(湖南大学) Shanghai AI Lab(上海人工智能实验室) University of Hildesheim(希尔德斯海姆大学)

AI总结 本文提出HyProMeta框架,通过双曲元学习和可学习提示缓解标签噪声,提升开放集领域泛化的性能。

Comments Accepted to International Journal of Computer Vision (IJCV). The source code of this work is released at https://github.com/KPeng9510/HyProMeta

URL PDF HTML 收藏
2406.02978 2025-12-29 cs.CV

Self-Supervised Skeleton-Based Action Representation Learning: A Benchmark and Beyond

自监督骨骼基动作表示学习:一个基准与更远

Jiahang Zhang, Lilang Lin, Shuai Yang, Jiaying Liu

机构 * Wangxuan Institute of Computer Technology(王轩计算机技术研究所) Peking University(北京大学)

AI总结 本文提出了一种新的自监督学习方法,通过整合多粒度的表示学习目标,提升骨骼动作表示的泛化能力,并在多个下游任务中验证了其有效性。

Comments IJCV 2025

URL PDF HTML 收藏
2512.19725 2025-12-24 cs.LG

Out-of-Distribution Detection for Continual Learning: Design Principles and Benchmarking

持续学习中的分布外检测:设计原则与基准测试

Srishti Gupta, Riccardo Balia, Daniele Angioni, Fabio Brau, Maura Pintor, Ambra Demontis, Alessandro Sebastian, Salvatore Mario Carta, Fabio Roli, Battista Biggio

AI总结 本文探讨了持续学习中分布外检测的设计原则与基准测试,旨在提升AI系统在动态环境中的适应性和鲁棒性。

Comments International Journal of Computer Vision

URL PDF HTML 收藏
2207.09775 2025-12-23 cs.CV

Rethinking Open-Set Object Detection: Issues, a New Formulation, and Taxonomy

重新思考开集目标检测:问题、新公式和分类

Yusuke Hosoya, Masanori Suganuma, Takayuki Okatani

机构 * RIKEN Center for AIP(理化学研究所AIP研究中心)

AI总结 本文重新思考开集目标检测问题,提出新的公式和分类,揭示现有方法在检测未知对象时的误分类问题。

Comments Accepted to IJCV

URL PDF HTML 收藏
2512.16113 2025-12-19 cs.CV

Flexible Camera Calibration using a Collimator System

基于校准系统的灵活相机校准

Shunkun Liang, Banglei Guan, Zhenbao Yu, Dongcai Tan, Pengju Sun, Zibin Liu, Qifeng Yu, Yang Shang

AI总结 本文提出了一种基于校准系统的灵活相机校准方法,通过引入角度不变约束和球面运动模型,实现了无需相机运动的快速校准。

Journal ref Liang S, Guan B, Yu Z, et al. Flexible Camera Calibration using a Collimator System[J]. International Journal of Computer Vision, 2025, 133(11): 8127-8150

URL PDF HTML 收藏
2402.02242 2025-12-10 cs.CV cs.LG

Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey and Benchmark

预训练视觉模型的参数高效微调:调查与基准

Yi Xin, Jianjiang Yang, Siqi Luo, Yuntao Du, Qi Qin, Kangrui Cen, Yangfan He, Zhiwei Zhang, Bin Fu, Xiaokang Yang, Guangtao Zhai, Ming-Hsuan Yang, Xiaohong Liu

机构 * Nanjing University(南京大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Jiao Tong University(上海交通大学) University of Bristol(布里斯托大学) Shandong University(山东大学) University of Sydney(悉尼大学) University of Minnesota Twin Cities(明尼苏达大学双城分校) The Pennsylvania State University(宾夕法尼亚州立大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) University of California at Merced(加州大学默塞德分校)

AI总结 本文调查了预训练视觉模型参数高效微调的最新进展,提出了V-PEFT Bench基准,并探讨了未来研究方向。

Comments Submitted to IJCV

URL PDF HTML 收藏
2308.09388 2025-12-09 cs.CV

Diffusion Models for Image Restoration and Enhancement: A Comprehensive Survey

扩散模型在图像修复与增强中的应用:全面综述

Xin Li, Yulin Ren, Xin Jin, Cuiling Lan, Xingrui Wang, Wenjun Zeng, Xinchao Wang, Zhibo Chen

机构 * University of Science and Technology of China(科学技术大学) National University of Singapore(国立新加坡大学) Eastern Institute for Advanced Study(东部高级研究机构) Microsoft Research Asia(微软亚洲研究院)

AI总结 本文首次全面综述了基于扩散模型的图像修复与增强方法,涵盖学习范式、条件策略、框架设计等,并提出未来研究方向。

Comments Accepted by IJCV 2025

URL PDF HTML 收藏
2511.18983 2025-11-25 cs.CV

UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake Detection

UMCL: 单模生成多模对比学习用于跨压缩率深度伪造检测

Ching-Yi Lai, Chih-Yu Jian, Pei-Cheng Chuang, Chia-Ming Lee, Chih-Chung Hsu, Chiou-Ting Hsu, Chia-Wen Lin

AI总结 UMCL通过单模生成多模对比学习,提升跨压缩率深度伪造检测的鲁棒性和准确性。

Comments 24-page manuscript accepted to IJCV

URL PDF HTML 收藏
2511.14279 2025-11-19 cs.CV

Free Lunch to Meet the Gap: Intermediate Domain Reconstruction for Cross-Domain Few-Shot Learning

Tong Zhang, Yifan Zhao, Liangyu Wang, Jia Li

机构 * State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室) School of Computer Science and Engineering(计算机科学与工程学院) Beihang University(北航大学)

Comments Accepted to IJCV 2025

URL PDF HTML 收藏
2405.15239 2025-11-18 cs.CV

Brain3D: Generating 3D Objects from fMRI

Yuankun Yang, Li Zhang, Ziyang Xie, Zhiyuan Yuan, Jianfeng Feng, Xiatian Zhu, Yu-Gang Jiang

机构 * Fudan University(复旦大学) University of California, Los Angeles(加州大学洛杉矶分校) University of Surrey(萨里大学)

Comments IJCV 2025

URL PDF HTML 收藏
2511.11639 2025-11-18 cs.RO cs.CV

Image-based Morphological Characterization of Filamentous Biological Structures with Non-constant Curvature Shape Feature

Jie Fan, Francesco Visentin, Barbara Mazzolai, Emanuela Del Dottore

Comments This manuscript is a preprint version of the article currently under peer review at International Journal of Computer Vision (IJCV)

URL PDF HTML 收藏
2511.11009 2025-11-17 cs.LG cs.CV

Unsupervised Robust Domain Adaptation: Paradigm, Theory and Algorithm

Fuxiang Huang, Xiaowei Fu, Shiyu Ye, Lina Ma, Wen Li, Xinbo Gao, David Zhang, Lei Zhang

机构 * Chongqing Key Laboratory of Bio-perception and Multimodal Intelligent Information Processing(重庆生物感知与多模态智能信息处理重点实验室) Chongqing University(重庆大学) School of Microelectronics and Communication Engineering(微电子与通信工程学院) University of Electronic Science and Technology of China(电子科学与技术大学) Chongqing Key Laboratory of Image Cognition(重庆图像认知重点实验室) Chongqing University of Posts and Telecommunications(重庆邮电大学) School of Data Science(数据科学学院) Lingnan University Hong Kong(香港岭南大学) School of Science and Engineering(科学与工程学院) Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳))

Comments To appear in IJCV

URL PDF HTML 收藏
2411.17040 2025-10-14 cs.CV

Multimodal Alignment and Fusion: A Survey

Songtao Li, Hao Tang

机构 * Peking University(北京大学) Northeastern University(东北大学) Sydney Smart Technology College(悉尼智能技术学院) School of Computer Science, Peking University(北京大学计算机学院) The State Key Laboratory of Multimedia Information Processing(多媒体信息处理国家重点实验室)

Comments Accepted to IJCV 2025

URL PDF HTML 收藏
2509.22063 2025-09-29 cs.CV cs.SD

High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling

Chao Huang, Susan Liang, Yapeng Tian, Anurag Kumar, Chenliang Xu

机构 * Meta Reality Labs Research(Meta现实实验室)

Comments Accepted to IJCV

URL PDF HTML 收藏
2509.17638 2025-09-23 cs.CV cs.AI

A$^2$M$^2$-Net: Adaptively Aligned Multi-Scale Moment for Few-Shot Action Recognition

Zilin Gao, Qilong Wang, Bingbing Zhang, Qinghua Hu, Peihua Li

机构 * School of Information and Communication Engineering(信息与通信工程学院) Dalian University of Technology(大连理工大学) College of Intelligence and Computing(智能与计算学院) Tianjin University(天津大学) Haihe Laboratory of Information Technology Application Innovation(信息科技应用创新海河实验室) School of Computer Science and Engineering(计算机科学与工程学院) Dalian Minzu University(大连民族大学)

Comments 27 pages, 13 figures, 7 tables

Journal ref Published in IJCV, 2025

URL PDF HTML 收藏
2509.09962 2025-09-15 cs.CV

An HMM-based framework for identity-aware long-term multi-object tracking from sparse and uncertain identification: use case on long-term tracking in livestock

Anne Marthe Sophie Ngo Bibinbe, Chiron Bang, Patrick Gagnon, Jamie Ahloy-Dallaire, Eric R. Paquet

机构 * Centre de développement du porc du Québec(魁北克猪发展中心)

Comments 13 pages, 7 figures, 1 table, accepted at CVPR animal workshop 2024, submitted to IJCV

URL PDF HTML 收藏
2312.06660 2025-09-09 cs.CV

EdgeSAM: Prompt-In-the-Loop Distillation for SAM

Chong Zhou, Xiangtai Li, Chen Change Loy, Bo Dai

机构 * Nanyang Technological University, Singapore(南洋理工大学) The University of Hong Kong, China(香港大学)

Comments IJCV 2025. Project page: https://www.mmlab-ntu.com/project/edgesam

URL PDF HTML 收藏
2509.05604 2025-09-09 cs.CV cs.AI

Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization

Jungin Park, Jiyoung Lee, Kwanghoon Sohn

机构 * Yonsei University(延世大学) Ewha Womans University(成均馆大学) Korea Institute of Science and Technology (KIST)(韩国科学技术院)

Comments Accepted to IJCV, 29 pages, 14 figures, 11 tables

URL PDF HTML 收藏
2301.05711 2025-09-03 cs.CV

OA-DET3D: Embedding Object Awareness as a General Plug-in for Multi-Camera 3D Object Detection

Xiaomeng Chu, Jiajun Deng, Jianmin Ji, Yu Zhang, Houqiang Li, Yanyong Zhang

机构 * University of Science and Technology of China(科学技术大学) The University of Sydney(悉尼大学)

Comments Accepted by IJCV after more than two years of reviewing. Original title: OA-BEV: Bringing Object Awareness to Bird's-Eye-View Representation for Multi-Camera 3D Object Detection

URL PDF HTML 收藏
2412.08321 2025-08-29 eess.SY cs.CV cs.SY

TGOSPA Metric Parameters Selection and Evaluation for Visual Multi-object Tracking

Jan Krejčí, Oliver Kost, Ondřej Straka, Yuxuan Xia, Lennart Svensson, Ángel F. García-Fernández

机构 * Department of Cybernetics, University of West Bohemia in Pilsen(西波希米亚大学控制系) Department of Automation, Shanghai Jiaotong University(上海交通大学自动化系) Signal Processing Group, Chalmers University of Technology(查尔姆斯理工大学信号处理组)

Comments Submitted to Springer International Journal of Computer Vision

URL PDF HTML 收藏
2508.17232 2025-08-27 cs.LG cs.CV stat.ML

Curvature Learning for Generalization of Hyperbolic Neural Networks

Xiaomeng Fan, Yuwei Wu, Zhi Gao, Mehrtash Harandi, Yunde Jia

机构 * Beijing Laboratory of Intelligent Information Technology(北京智能信息科技实验室) School of Computer Science, Beijing Institute of Technology (BIT)(北京理工大学计算机学院) Guangdong Laboratory of Machine Perception and Intelligent Computing(广东机器感知与智能计算实验室) Shenzhen MSU-BIT University(深圳MSU-BIT大学) Department of Electrical and Computer Systems Eng.(电气与计算机系统工程系) Data61

Comments Accepted by International Journal of Computer Vision (IJCV)

URL PDF HTML 收藏
2412.01240 2025-08-27 cs.CV

Inspiring the Next Generation of Segment Anything Models: Comprehensively Evaluate SAM and SAM 2 with Diverse Prompts Towards Context-Dependent Concepts under Different Scenes

Xiaoqi Zhao, Youwei Pang, Shijie Chang, Yuan Zhao, Lihe Zhang, Chenyang Yu, Hanqi Liu, Jiaming Zuo, Jinsong Ouyang, Weisi Lin, Georges El Fakhri, Huchuan Lu, Xiaofeng Liu

机构 * Yale University, USA(耶鲁大学) Nanyang Technological University, Singapore(南洋理工大学) Dalian University of Technology, China(大连理工大学) X3000 Inspection Co., Ltd, China(X3000检测有限公司)

Comments Under submission to International Journal of Computer Vision (IJCV)

URL PDF HTML 收藏
2402.11057 2025-08-14 cs.CV

Are you Struggling? Dataset and Baselines for Struggle Determination in Assembly Videos

Shijia Feng, Michael Wray, Brian Sullivan, Youngkyoon Jang, Casimir Ludwig, Iain Gilchrist, Walterio Mayol-Cuevas

机构 * Amazon(亚马逊)

Comments Accepted by International Journal of Computer Vision (IJCV, 2025)

URL PDF HTML 收藏