arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Computer Vision · 会议 · Computer Vision

共收录 4770
2508.04211 2025-08-07 cs.CV

What Holds Back Open-Vocabulary Segmentation?

Josip Šarić, Ivan Martinović, Matej Kristan, Siniša Šegvić

机构 * Faculty of Computer and Information Science(计算机与信息科学学院) Faculty of Electrical Engineering and Computing(电子工程与计算学院)

Comments Accepted for publication at ICCV 25 Workshop: What is Next in Multimodal Foundation Models?

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04122 2025-08-07 cs.CV

Conditional Latent Diffusion Models for Zero-Shot Instance Segmentation

Maximilian Ulmer, Wout Boerdijk, Rudolph Triebel, Maximilian Durner

机构 * German Aerospace Center (DLR)(德国航空航天中心) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Technical University of Munich(慕尼黑技术大学)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02629 2025-08-07 cs.RO cs.AI cs.CL

HyCodePolicy: Hybrid Language Controllers for Multimodal Monitoring and Decision in Embodied Agents

Yibin Liu, Zhixuan Liang, Zanxin Chen, Tianxing Chen, Mengkang Hu, Wanxi Dong, Congsheng Xu, Zhaoming Han, Yusen Qin, Yao Mu

Comments Accepted to ICCV 2025 Workshop on Multi-Modal Reasoning for Agentic Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19264 2025-08-07 cs.CV

SimMLM: A Simple Framework for Multi-modal Learning with Missing Modality

Sijie Li, Chen Chen, Jungong Han

机构 * School of Computer Science, University of Sheffield, UK(计算机科学学院,谢菲尔德大学)

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21884 2025-08-07 eess.IV cs.AI cs.CV cs.LG eess.SP

UnMix-NeRF: Spectral Unmixing Meets Neural Radiance Fields

Fabian Perez, Sara Rojas, Carlos Hinojosa, Hoover Rueda-Chacón, Bernard Ghanem

机构 * Universidad Industrial de Santander(圣安德烈大学) KAUST(王国立人工智能科技大学)

Comments Paper accepted at ICCV 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03123 2025-08-07 cs.CV

Dual-Expert Consistency Model for Efficient and High-Quality Video Generation

Zhengyao Lv, Chenyang Si, Tianlin Pan, Zhaoxi Chen, Kwan-Yee K. Wong, Yu Qiao, Ziwei Liu

机构 * Nanjing University(南京大学) The University of Hong Kong(香港大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Chinese Academy of Sciences(中国科学院大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室)

Comments This paper has been accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10905 2025-08-07 cs.AI cs.CV cs.LG

Learning to Inference Adaptively for Multimodal Large Language Models

Zhuoyan Xu, Khoi Duc Nguyen, Preeti Mukherjee, Saurabh Bagchi, Somali Chaterji, Yingyu Liang, Yin Li

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Purdue University(普渡大学) The University of Hong Kong(香港大学)

Comments Published at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20879 2025-08-07 cs.CV

egoPPG: Heart Rate Estimation from Eye-Tracking Cameras in Egocentric Systems to Benefit Downstream Vision Tasks

Björn Braun, Rayan Armani, Manuel Meier, Max Moebus, Christian Holz

机构 * Department of Computer Science, ETH Zürich(计算机科学系,苏黎世联邦理工学院)

Comments Accepted for publication at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04449 2025-08-07 cs.CV cs.CL

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay

Jun Zhang, Desen Meng, Zhengming Zhang, Zhenpeng Huang, Tao Wu, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) China Mobile Research Institute(中国移动研究院) Shanghai AI Lab(上海AI实验室)

Comments Accepted by ICCV 2025; Code released at https://github.com/MCG-NJU/p-MoD

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03859 2025-08-07 cs.CV

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Hui Zhang, Dexiang Hong, Yitong Wang, Jie Shao, Xinglong Wu, Zuxuan Wu, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究院) Fudan University(复旦大学) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Bytedance Intelligent Creation(字节跳动智能创作)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01812 2025-08-07 cs.CV

V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and Prediction

Zewei Zhou, Hao Xiang, Zhaoliang Zheng, Seth Z. Zhao, Mingyue Lei, Yun Zhang, Tianhui Cai, Xinyi Liu, Johnson Liu, Maheswari Bajji, Xin Xia, Zhiyu Huang, Bolei Zhou, Jiaqi Ma

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

Comments ICCV 2025, Website link: https://mobility-lab.seas.ucla.edu/v2xpnp/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03695 2025-08-06 cs.CV

Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition

Pulkit Kumar, Shuaiyi Huang, Matthew Walmer, Sai Saketh Rambhatla, Abhinav Shrivastava

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

Comments Accepted at ICCV 2025; First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03596 2025-08-06 cs.CV cs.AI

MetaScope: Optics-Driven Neural Network for Ultra-Micro Metalens Endoscopy

Wuyang Li, Wentao Pan, Xiaoyuan Liu, Zhendong Luo, Chenxin Li, Hengyu Liu, Din Ping Tsai, Mu Ku Chen, Yixuan Yuan

机构 * The Chinese University of Hong Kong(香港中文大学) City University of Hong Kong(香港城市大学) EPFL(苏黎世联邦理工学院)

Comments ICCV 2025 (Highlight); Project Page: https://cuhk-aim-group.github.io/MetaScope/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03481 2025-08-06 cs.CV cs.AI cs.CL

Draw Your Mind: Personalized Generation via Condition-Level Modeling in Text-to-Image Diffusion Models

Hyungjin Kim, Seokho Ahn, Young-Duk Seo

机构 * Department of Electrical and Computer Engineering, Inha University(电子与计算机工程系,inha大学)

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17748 2025-08-06 cs.LG cs.AI cs.CV stat.ML

Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and Compressibility

Melih Barsbey, Lucas Prieto, Stefanos Zafeiriou, Tolga Birdal

机构 * Imperial College London(伦敦帝国学院)

Comments Accepted at ICCV 2025, 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03402 2025-08-06 cs.CV cs.AI cs.LG

SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models

Pingchuan Ma, Xiaopei Yang, Yusong Li, Ming Gui, Felix Krause, Johannes Schusterbauer, Björn Ommer

机构 * Munich Center for Machine Learning (MCML)(慕尼黑機器學習中心)

Comments ICCV 2025, Project Page: https://compvis.github.io/SCFlow/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03164 2025-08-06 cs.CV cs.AI cs.CL

ChartCap: Mitigating Hallucination of Dense Chart Captioning

Junyoung Lim, Jaewoo Ahn, Gunhee Kim

机构 * Seoul National University(首尔国立大学)

Comments ICCV 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03118 2025-08-06 cs.CV

H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction

Heng Jia, Linchao Zhu, Na Zhao

机构 * ReLER Lab, CCAI, Zhejiang University(浙江大学ReLER实验室,中国计算机协会,浙江大学) The State Key Lab of Brain-Machine Intelligence, Zhejiang University(浙江大学脑机智能国家重点实验室) Singapore University of Technology and Design(新加坡科技设计大学)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02987 2025-08-06 cs.CV

Adversarial Attention Perturbations for Large Object Detection Transformers

Zachary Yahn, Selim Furkan Tekin, Fatih Ilhan, Sihao Hu, Tiansheng Huang, Yichang Xu, Margaret Loper, Ling Liu

机构 * Georgia Institute of Technology(佐治亚理工学院) Georgia Tech Research Institute(佐治亚理工研究 institute)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02936 2025-08-06 cs.AI

AQUAH: Automatic Quantification and Unified Agent in Hydrology

Songkun Yan, Zhi Li, Siyu Zhu, Yixin Wen, Mofan Zhang, Mengye Chen, Jie Cao, Yang Hong

Comments 8 pages, 5 figures, 2025 ICCV SEA workshop paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02905 2025-08-06 cs.CV cs.SD eess.AS

How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes

Mahnoor Fatima Saad, Ziad Al-Halah

机构 * University of Utah(犹他大学)

Comments Accepted to ICCV 2025. Project Page: https://mahnoor-fatima-saad.github.io/m-capa.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01242 2025-08-06 cs.GR cs.CV

MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh

Shuangkang Fang, I-Chao Shen, Yufeng Wang, Yi-Hsuan Tsai, Yi Yang, Shuchang Zhou, Wenrui Ding, Takeo Igarashi, Ming-Hsuan Yang

机构 * Beihang University(北航大学) The University of Tokyo(东京大学) Atmanity Inc.(Atmanity公司) StepFun Inc.(StepFun公司) UC Merced(加州大学默塞德分校)

Comments Accepted by ICCV. Project Website: https://sk-fun.fun/MeshLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12137 2025-08-06 cs.CV

AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous Driving

Jiawei Xu, Kai Deng, Zexin Fan, Shenlong Wang, Jin Xie, Jian Yang

机构 * College of Computer Science, Nankai University(南开大学计算机科学学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11986 2025-08-06 cs.CV

Style Composition within Distinct LoRA modules for Traditional Art

Jaehyun Lee, Wonhark Park, Wonsik Shin, Hyunho Lee, Hyoung Min Na, Nojun Kwak

机构 * Seoul National University(首尔国立大学) Kyunghee University(庆熙大学)

Comments Accepted to ICCV 2025 Workshop(WCCA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17213 2025-08-06 cs.CV cs.AI cs.RO

Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario Generation

Xiuyu Yang, Shuhan Tan, Philipp Krähenbühl

机构 * UT Austin(德克萨斯大学奥斯汀分校)

Comments ICCV 2025. Project page: https://orangesodahub.github.io/InfGen Code: https://github.com/OrangeSodahub/infgen

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05787 2025-08-06 cs.CV

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs

Ivan Rodin, Tz-Ying Wu, Kyle Min, Sharath Nittur Sridhar, Antonino Furnari, Subarna Tripathi, Giovanni Maria Farinella

机构 * University of Catania(卡塔尼亚大学) Intel Labs(英特尔实验室)

Comments Accepted to SAUAFG Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05307 2025-08-06 cs.CV

PRE-Mamba: A 4D State Space Model for Ultra-High-Frequent Event Camera Deraining

Ciyu Ruan, Ruishan Guo, Zihang Gong, Jingao Xu, Wenhan Yang, Xinlei Chen

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Harbin Institute of Technology(哈尔滨工业大学) Carnegie Mellon University(卡内基梅隆大学) Pengcheng Laboratory(鹏城实验室)

Comments This version is the camera-ready version accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00482 2025-08-06 cs.CV cs.AI

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers

Kwon Byung-Ki, Qi Dai, Lee Hyoseok, Chong Luo, Tae-Hyun Oh

机构 * POSTECH Microsoft Research Asia(微软亚洲研究院) KAIST(韩国科学技术院)

Comments Accepted to IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Project page: https://byungki-k.github.io/JointDiT/ Code: https://github.com/kaist-ami/JointDiT

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01647 2025-08-06 cs.CV

FlowR: Flowing from Sparse to Dense 3D Reconstructions

Tobias Fischer, Samuel Rota Bulò, Yung-Hsu Yang, Nikhil Keetha, Lorenzo Porzi, Norman Müller, Katja Schwarz, Jonathon Luiten, Marc Pollefeys, Peter Kontschieder

机构 * ETH Zürich(苏黎世联邦理工学院) Meta Reality Labs Zürich(Meta现实实验室(苏黎世)) CMU(卡内基梅隆大学)

Comments ICCV 2025 Highlight. Project page is available at https://tobiasfshr.github.io/pub/flowr

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17350 2025-08-06 cs.CV

Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer

Qingyu Shi, Jianzong Wu, Jinbin Bai, Jiangning Zhang, Lu Qi, Yunhai Tong, Xiangtai Li

机构 * PKU(北京大学) NTU(国立新加坡大学) NUS(新加坡国立大学) ZJU(浙江大学) UC Merced(加州大学默塞德分校)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏