arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Computer Vision · 会议 · Computer Vision

2025-08-07 至 2025-08-07 共收录 17
2508.04625 2025-08-07 cs.CV cs.CE

FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging

Zichen Tang, Haihong E, Jiacheng Liu, Zhongjun Yang, Rongjin Li, Zihua Rong, Haoyang He, Zhuodi Hao, Xinyang Hu, Kun Ji, Ziyan Ma, Mengyuan Ji, Jun Zhang, Chenghao Ma, Qianhe Zheng, Yang Liu, Yiling Huang, Xinyi Hu, Qing Huang, Zijian Xie, Shiyao Peng

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

Comments Accepted by ICCV 2025. arXiv admin note: text overlap with arXiv:2311.06602 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04681 2025-08-07 cs.CV

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

Liang Xu, Chengqun Yang, Zili Lin, Fei Xu, Yifan Liu, Congsheng Xu, Yiyi Zhang, Jie Qin, Xingdong Sheng, Yunhui Liu, Xin Jin, Yichao Yan, Wenjun Zeng, Xiaokang Yang

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能教育部重点实验室,上海交通大学AI研究院) Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波数字孪生研究所,东部技术研究所) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) MoE Key Lab of AI, School of Computer Science, Shanghai Jiao Tong University(人工智能教育部重点实验室,上海交通大学计算机学院) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Lenovo(联想公司)

Comments Accepted to ICCV 2025. Project Page: https://liangxuy.github.io/InterVLA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04659 2025-08-07 cs.CV

PixCuboid: Room Layout Estimation from Multi-view Featuremetric Alignment

Gustav Hanning, Kalle Åström, Viktor Larsson

机构 * Lund University(隆德大学)

Comments Accepted at the ICCV 2025 Workshop on Large Scale Cross Device Localization

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04648 2025-08-07 astro-ph.IM cs.CV

Super Resolved Imaging with Adaptive Optics

Robin Swanson, Esther Y. H. Lin, Masen Lamb, Suresh Sivanandam, Kiriakos N. Kutulakos

机构 * University of Toronto(多伦多大学) Dunlap Institute for Astronomy & Astrophysics(天文学与天体物理学敦洛普研究所) International Gemini Observatory(国际Gemini天文台) University of Victoria(维多利亚大学)

Comments Accepted to ICCV 2025 (IEEE/CVF International Conference on Computer Vision)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04642 2025-08-07 cs.RO cs.CV

RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case

Baihui Xiao, Chengjian Feng, Zhijian Huang, Feng yan, Yujie Zhong, Lin Ma

机构 * Meituan(美团) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04546 2025-08-07 cs.CV

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

Minghang Zheng, Yuxin Peng, Benyuan Sun, Yi Yang, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王宣计算机技术研究所) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) Central Media Technology Institute, Huawei(华为中央媒体技术研究所)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04211 2025-08-07 cs.CV

What Holds Back Open-Vocabulary Segmentation?

Josip Šarić, Ivan Martinović, Matej Kristan, Siniša Šegvić

机构 * Faculty of Computer and Information Science(计算机与信息科学学院) Faculty of Electrical Engineering and Computing(电子工程与计算学院)

Comments Accepted for publication at ICCV 25 Workshop: What is Next in Multimodal Foundation Models?

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04122 2025-08-07 cs.CV

Conditional Latent Diffusion Models for Zero-Shot Instance Segmentation

Maximilian Ulmer, Wout Boerdijk, Rudolph Triebel, Maximilian Durner

机构 * German Aerospace Center (DLR)(德国航空航天中心) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Technical University of Munich(慕尼黑技术大学)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02629 2025-08-07 cs.RO cs.AI cs.CL

HyCodePolicy: Hybrid Language Controllers for Multimodal Monitoring and Decision in Embodied Agents

Yibin Liu, Zhixuan Liang, Zanxin Chen, Tianxing Chen, Mengkang Hu, Wanxi Dong, Congsheng Xu, Zhaoming Han, Yusen Qin, Yao Mu

Comments Accepted to ICCV 2025 Workshop on Multi-Modal Reasoning for Agentic Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19264 2025-08-07 cs.CV

SimMLM: A Simple Framework for Multi-modal Learning with Missing Modality

Sijie Li, Chen Chen, Jungong Han

机构 * School of Computer Science, University of Sheffield, UK(计算机科学学院,谢菲尔德大学)

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21884 2025-08-07 eess.IV cs.AI cs.CV cs.LG eess.SP

UnMix-NeRF: Spectral Unmixing Meets Neural Radiance Fields

Fabian Perez, Sara Rojas, Carlos Hinojosa, Hoover Rueda-Chacón, Bernard Ghanem

机构 * Universidad Industrial de Santander(圣安德烈大学) KAUST(王国立人工智能科技大学)

Comments Paper accepted at ICCV 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03123 2025-08-07 cs.CV

Dual-Expert Consistency Model for Efficient and High-Quality Video Generation

Zhengyao Lv, Chenyang Si, Tianlin Pan, Zhaoxi Chen, Kwan-Yee K. Wong, Yu Qiao, Ziwei Liu

机构 * Nanjing University(南京大学) The University of Hong Kong(香港大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Chinese Academy of Sciences(中国科学院大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室)

Comments This paper has been accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10905 2025-08-07 cs.AI cs.CV cs.LG

Learning to Inference Adaptively for Multimodal Large Language Models

Zhuoyan Xu, Khoi Duc Nguyen, Preeti Mukherjee, Saurabh Bagchi, Somali Chaterji, Yingyu Liang, Yin Li

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Purdue University(普渡大学) The University of Hong Kong(香港大学)

Comments Published at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20879 2025-08-07 cs.CV

egoPPG: Heart Rate Estimation from Eye-Tracking Cameras in Egocentric Systems to Benefit Downstream Vision Tasks

Björn Braun, Rayan Armani, Manuel Meier, Max Moebus, Christian Holz

机构 * Department of Computer Science, ETH Zürich(计算机科学系,苏黎世联邦理工学院)

Comments Accepted for publication at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04449 2025-08-07 cs.CV cs.CL

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay

Jun Zhang, Desen Meng, Zhengming Zhang, Zhenpeng Huang, Tao Wu, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) China Mobile Research Institute(中国移动研究院) Shanghai AI Lab(上海AI实验室)

Comments Accepted by ICCV 2025; Code released at https://github.com/MCG-NJU/p-MoD

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03859 2025-08-07 cs.CV

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Hui Zhang, Dexiang Hong, Yitong Wang, Jie Shao, Xinglong Wu, Zuxuan Wu, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究院) Fudan University(复旦大学) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Bytedance Intelligent Creation(字节跳动智能创作)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01812 2025-08-07 cs.CV

V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and Prediction

Zewei Zhou, Hao Xiang, Zhaoliang Zheng, Seth Z. Zhao, Mingyue Lei, Yun Zhang, Tianhui Cai, Xinyi Liu, Johnson Liu, Maheswari Bajji, Xin Xia, Zhiyu Huang, Bolei Zhou, Jiaqi Ma

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

Comments ICCV 2025, Website link: https://mobility-lab.seas.ucla.edu/v2xpnp/

详情

展开后加载摘要…

URL PDF HTML 收藏