arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Computer Vision · 会议 · Computer Vision

共收录 4770
2507.15569 2025-07-22 cs.CV

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

Xiaoyi Bao, Chenwei Xie, Hao Tang, Tingyu Weng, Xiaofeng Wang, Yun Zheng, Xingang Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Alibaba Group(阿里巴巴集团) Peking University(北京大学) Luoyang Institute for Robot and Intelligent Equipment(洛阳机器人与智能装备研究所)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15454 2025-07-22 cs.GR cs.AI cs.CV cs.HC

ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting

Ruijie Zhu, Mulin Yu, Linning Xu, Lihan Jiang, Yixuan Li, Tianzhu Zhang, Jiangmiao Pang, Bo Dai

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15399 2025-07-22 cs.GR cs.CV

Blended Point Cloud Diffusion for Localized Text-guided Shape Editing

Etai Sella, Noam Atia, Ron Mokady, Hadar Averbuch-Elor

Comments Accepted to ICCV 2025. Project Page: https://tau-vailab.github.io/BlendedPC/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15381 2025-07-22 cs.LG cs.AI cs.CV

To Label or Not to Label: PALM -- A Predictive Model for Evaluating Sample Efficiency in Active Learning Models

Julia Machnio, Mads Nielsen, Mostafa Mehdipour Ghazi

机构 * Pioneer Centre for AI(先锋人工智能中心) University of Copenhagen(哥本哈根大学)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15365 2025-07-22 cs.CV

DAViD: Data-efficient and Accurate Vision Models from Synthetic Data

Fatemeh Saleh, Sadegh Aliakbarian, Charlie Hewitt, Lohit Petikam, Xiao-Xian, Antonio Criminisi, Thomas J. Cashman, Tadas Baltrušaitis

机构 * Microsoft, Cambridge, UK(微软公司,剑桥,英国)

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15260 2025-07-22 cs.LG

CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE Solvers

Jiaqi Han, Haotian Ye, Puheng Li, Minkai Xu, James Zou, Stefano Ermon

机构 * Stanford University(斯坦福大学)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15257 2025-07-22 cs.CV

MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnP

Pei An, Jiaqi Yang, Muyao Peng, You Yang, Qiong Liu, Xiaolin Wu, Liangliang Nan

机构 * Huazhong University of Science and Technology, China(华中科技大学) Northwestern Polytechnical University, China(西北工业大学) McMaster University, Canada(麦马士威大学) Delft University of Technology, Netherlands(代尔夫特理工大学)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15249 2025-07-22 cs.CV

FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers

Yanbing Zhang, Zhe Wang, Qin Zhou, Mengping Yang

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15028 2025-07-22 cs.CV

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

Yuanhan Zhang, Yunice Chew, Yuhao Dong, Aria Leo, Bo Hu, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室)

Comments ICCV 2025; Project page: https://zhangyuanhan-ai.github.io/video-tt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14935 2025-07-22 cs.CV

Open-set Cross Modal Generalization via Multimodal Unified Representation

Hai Huang, Yan Xia, Shulei Wang, Hanting Wang, Minghui Fang, Shengpeng Ji, Sashuai Zhou, Tao Jin, Zhou Zhao

机构 * Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08416 2025-07-22 cs.CV

InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes

Zesong Yang, Bangbang Yang, Wenqi Dong, Chenxuan Cao, Liyuan Cui, Yuewen Ma, Zhaopeng Cui, Hujun Bao

机构 * State Key Lab of CAD & CG(CAD与CG国家重点实验室) ByteDance(字节跳动)

Comments Accepted by ICCV 2025. Project page: https://zju3dv.github.io/instascene/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07778 2025-07-22 cs.LG cs.AI cs.CV

Synchronizing Task Behavior: Aligning Multiple Tasks during Test-Time Training

Wooseong Jeong, Jegyeong Cho, Youngho Yoon, Kuk-Jin Yoon

机构 * Visual Intelligence Lab., KAIST, Korea(韩国釜山科学技术院视觉智能实验室)

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07485 2025-07-22 cs.LG cs.AI cs.CV

Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task Learning

Wooseong Jeong, Kuk-Jin Yoon

机构 * KAIST(韩国科学技术院)

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06060 2025-07-22 cs.CV cs.AI

VisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis

Alexandre Symeonidis-Herzig, Özge Mercanoğlu Sincan, Richard Bowden

机构 * CVSSP, University of Surrey, United Kingdom(CVSSP,英国萨里大学)

Comments Accepted in International Conference on Computer Vision (ICCV) Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02939 2025-07-22 cs.LG cs.AI cs.CV

Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting

Yuqi Li, Chuanguang Yang, Hansheng Zeng, Zeyu Dong, Zhulin An, Yongjun Xu, Yingli Tian, Hao Wu

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) The University of Hong Kong(香港大学) The City University of New York(纽约城市大学) Tsinghua University(清华大学)

Comments Accepted by ICCV-2025, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21313 2025-07-22 cs.CV

HORT: Monocular Hand-held Objects Reconstruction with Transformers

Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, Cordelia Schmid

机构 * Inria, École normale supérieure, CNRS, PSL Research University(法国国家科学研究中心、巴黎高等师范学院、PSL研究大学) Imperial College London(伦敦帝国理工学院)

Comments Accepted by ICCV 2025. Project Page: https://zerchen.github.io/projects/hort.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19457 2025-07-22 cs.CV cs.RO

G-DexGrasp: Generalizable Dexterous Grasping Synthesis Via Part-Aware Prior Retrieval and Prior-Assisted Generation

Juntao Jian, Xiuping Liu, Zixuan Chen, Manyi Li, Jian Liu, Ruizhen Hu

机构 * Shenzhen University(深圳大学) Dalian University of Technology(大连理工大学) Shandong University(山东大学) Shenyang University of Technology(沈阳理工大学)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12897 2025-07-22 cs.LG cs.AI

Federated Continual Instruction Tuning

Haiyang Guo, Fanhu Zeng, Fei Zhu, Wenzhuo Liu, Da-Han Wang, Jian Xu, Xu-Yao Zhang, Cheng-Lin Liu

机构 * School of Advanced Interdisciplinary Sciences, UCAS(中国科学院大学先进交叉学科学院) MAIS, CASIA(中国科学院自动化所人工智能研究所) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Centre for Artificial Intelligence and Robotics, HKISI-CAS(香港科技大学人工智能与机器人中心,中国科学院)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06505 2025-07-22 cs.CV cs.AI

DynamicID: Zero-Shot Multi-ID Image Personalization with Flexible Facial Editability

Xirui Hu, Jiahao Wang, Hao Chen, Weizhan Zhang, Benqi Wang, Yikun Li, Haishun Nan

机构 * Xi’an Jiaotong University(西安交通大学) Film AI Lab, Western Movie Group(西方电影集团影视人工智能实验室)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06273 2025-07-22 cs.CV cs.MM cs.SD eess.AS

Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

Jeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis, Yong Man Ro

机构 * KAIST(韩国科学技术院) Imperial College London(伦敦帝国理工学院)

Comments Accepted at ICCV 2025. Code available at: https://github.com/JeongHun0716/zero-avsr

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01425 2025-07-22 cs.CV

Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation

Xincheng Shuai, Henghui Ding, Zhenyuan Qin, Hao Luo, Xingjun Ma, Dacheng Tao

机构 * Fudan University(复旦大学) DAMO Academy, Alibaba group(阿里集团 DAMO 院) Hupan Lab(虎派实验室) Nanyang Technological University, Singapore(新加坡南洋理工大学)

Comments ICCV 2025, Project Page: https://henghuiding.github.io/SynFMC/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17769 2025-07-22 cs.CV

Omegance: A Single Parameter for Various Granularities in Diffusion-Based Synthesis

Xinyu Hou, Zongsheng Yue, Xiaoming Li, Chen Change Loy

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室)

Comments ICCV 2025 Camera Ready. Project page: https://itsmag11.github.io/Omegance/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10440 2025-07-22 cs.CV

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Guowei Xu, Peng Jin, Ziang Wu, Hao Li, Yibing Song, Lichao Sun, Li Yuan

Comments 17 pages, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10360 2025-07-22 cs.CV cs.AI

HaSPeR: An Image Repository for Hand Shadow Puppet Recognition

Syed Rifat Raiyan, Zibran Zarif Amio, Sabbir Ahmed

机构 * Systems and Software Lab (SSL) of the Department of Computer Science and Engineering at the Islamic University of Technology(伊斯兰科技大学计算机科学与工程系系统与软件实验室) Networking Research Group of the Department of Computer Science and Engineering at the Islamic University of Technology(伊斯兰科技大学计算机科学与工程系网络研究组) Computer Vision Lab (CVLab) of the Department of Computer Science and Engineering at the Islamic University of Technology(伊斯兰科技大学计算机科学与工程系计算机视觉实验室) BrillMark LLC.(BrillMark LLC)

Comments Accepted in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops (WCCA 2025 Oral), 17 pages, 111 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14826 2025-07-22 cs.CV

PHATNet: A Physics-guided Haze Transfer Network for Domain-adaptive Real-world Image Dehazing

Fu-Jen Tsai, Yan-Tsung Peng, Yen-Yu Lin, Chia-Wen Lin

机构 * National Tsing Hua University(清华大学) MediaTek(联发科) National Chengchi University(台北市立成功大学) National Yang Ming Chiao Tung University(交通大学)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14797 2025-07-22 cs.CV

Distilling Parallel Gradients for Fast ODE Solvers of Diffusion Models

Beier Zhu, Ruoyu Wang, Tong Zhao, Hanwang Zhang, Chi Zhang

机构 * Nanyang Technological University(南洋理工大学) Westlake University(西湖大学)

Comments To appear in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14596 2025-07-22 cs.CV

DiSCO-3D : Discovering and segmenting Sub-Concepts from Open-vocabulary queries in NeRF

Doriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot, Loïc Barthe

机构 * Université Paris-Saclay, CEA, List(巴黎萨克雷大学,CEA,List) IRIT, Université de Toulouse, CNRS(IRIT,图卢兹大学,CNRS)

Comments Published at ICCV'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14505 2025-07-22 cs.CV

DCHM: Depth-Consistent Human Modeling for Multiview Detection

Jiahao Ma, Tianyu Wang, Miaomiao Liu, David Ahmedt-Aristizabal, Chuong Nguyen

机构 * Australian National University(澳大利亚国立大学) CSIRO Data61(CSIRO数据61)

Comments multi-view detection, sparse-view reconstruction

Journal ref ICCV`2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14449 2025-07-22 cs.CV

IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale Benchmark

Zhe Cao, Jin Zhang, Ruiheng Zhang

机构 * Beijing Institute of Technology(北京理工大学)

Comments 11 pages, 7 figures. This paper is accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12462 2025-07-22 cs.CV

SpatialTrackerV2: 3D Point Tracking Made Easy

Yuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev, Yuri Makarov, Bingyi Kang, Xing Zhu, Hujun Bao, Yujun Shen, Xiaowei Zhou

机构 * Zhejiang University(浙江大学) Oxford(牛津大学) Ant Group(蚂蚁集团) Pixelwise AI Bytedance Seed(字节跳动种子)

Comments International Conference on Computer Vision, ICCV 2025. Huggingface Demo: https://huggingface.co/spaces/Yuxihenry/SpatialTrackerV2, Code: https://github.com/henry123-boy/SpaTrackerV2

详情

展开后加载摘要…

URL PDF HTML 收藏