arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

International Conference on Computer Vision · 会议 · Computer Vision

共收录 4770
2509.23867 2025-11-25 cs.CV

Sim-DETR: Unlock DETR for Temporal Sentence Grounding

Sim-DETR: 解锁用于时间句子定位的DETR

Jiajin Tang, Zhengxuan Wei, Yuchen Zhu, Cheng Shi, Guanbin Li, Liang Lin, Sibei Yang

机构 * ShanghaiTech University(上海科技大学) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

AI总结 Sim-DETR通过改进DETR的解码器层,解决了时间句子定位中的查询冲突问题,提升了模型性能。

Comments This work is accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07608 2025-11-25 cs.CV

Faster and Better 3D Splatting via Group Training

更高效且更好的3D点云绘制方法:通过分组训练

Chengbo Wang, Guozheng Ma, Yifei Xue, Yizhen Lao

机构 * Hunan University(湖南大学) Nanyang Technological University(南洋理工大学) Lushan Innovation Lab(庐山创新实验室) Key Laboratory of Digital Culture Smart Design Technology, Minstry of Culture and Tourism(文化与旅游部数字文化智能设计技术重点实验室)

AI总结 本文提出分组训练方法,通过将高斯基本形体分组优化训练效率和渲染质量,实现更高效的3D点云绘制

Comments Accepted to ICCV 2025. Code is available at https://github.com/Chengbo-Wang/3DGS-with-Group-Training

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.09607 2025-11-25 cs.LG cs.RO

Description of Corner Cases in Automated Driving: Goals and Challenges

自动驾驶中角落情况的描述:目标与挑战

Daniel Bogdoll, Jasmin Breitenstein, Florian Heidecker, Maarten Bieshaar, Bernhard Sick, Tim Fingscheidt, J. Marius Zöllner

AI总结 本文探讨了自动驾驶中角落情况的描述目标与挑战,强调了机器可解释描述的重要性及当前研究的不足。

Comments Daniel Bogdoll, Jasmin Breitenstein and Florian Heidecker contributed equally. Accepted for publication at ICCV 2021 ERCVAD Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17888 2025-11-25 cs.CV

MINDiff: Mask-Integrated Negative Attention for Controlling Overfitting in Text-to-Image Personalization

MINDiff: 集成掩码的负关注用于控制文本到图像个性化中的过拟合

Seulgi Jeong, Jaeil Kim

机构 * Kyungpook National University(庆尚国立大学)

AI总结 MINDiff通过在推理过程中引入负关注机制,有效控制文本到图像个性化中的过拟合问题,无需修改模型架构即可提升文本对齐度和主题保真度。

Comments Accepted at ICCV 2025 Personalization in Generative AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17647 2025-11-25 cs.LG cs.AI

MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric Sequence

MamTiff-CAD: 多尺度潜在扩散与Mamba+用于复杂参数序列

Liyuan Deng, Yunpeng Bai, Yongkang Dai, Xiaoshui Huang, Hongping Gan, Dongshuo Huang, Hao jiacheng, Yilei Shi

机构 * Northwestern Polytechnical University(西北工业大学) National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学) Nanchang University(南昌大学)

AI总结 MamTiff-CAD通过结合Mamba+和Transformer的多尺度潜在扩散模型,有效生成复杂CAD参数序列,实现长序列生成任务的高性能表现。

Comments ICCV 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03931 2025-11-25 cs.CV

MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers

MagicMirror: 视频扩散变换器中的身份保留视频生成

Yuechen Zhang, Yaoyang Liu, Bin Xia, Bohao Peng, Zexin Yan, Eric Lo, Jiaya Jia

AI总结 MagicMirror通过双分支特征提取器和两阶段训练策略,在视频扩散变换器中实现高质量身份保留视频生成,优于现有方法。

Comments ICCV 2025, It is best viewed in Acrobat. Project Page: https://julianjuaner.github.io/projects/MagicMirror/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17065 2025-11-24 stat.ME

Shape Analysis of Euclidean Curves under Frenet-Serret Framework

基于弗朗谢-塞尔特框架的欧几里得曲线形状分析

Perrine Chassat, Juhyun Park, Nicolas Brunel

AI总结 本文提出基于弗朗谢-塞尔特框架的欧几里得曲线形状分析方法,通过平方根曲率变换构建黎曼几何,以更全面地捕捉曲线的几何特征,并应用于手语轨迹分析。

Journal ref IEEE/CVF International Conference on Computer Vision (ICCV) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12972 2025-11-24 cs.CV cs.AI

Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning

对齐视觉与语言:无需注释的多模态知识图谱构建以增强大语言模型推理

Junming Liu, Siyuan Meng, Yanting Gao, Song Mao, Pinlong Cai, Guohang Yan, Yirong Chen, Zilin Bian, Ding Wang, Botian Shi

机构 * Tongji University(同济大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) East China Normal University(华东师范大学) Stanford University(斯坦福大学) New York University(纽约大学)

AI总结 本文提出 VaLiK 方法,通过跨模态信息补充构建无需注释的多模态知识图谱,提升大语言模型推理能力。

Comments 14 pages, 7 figures, 6 tables; Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05456 2025-11-18 cs.CV

SITE: towards Spatial Intelligence Thorough Evaluation

Wenqi Wang, Reuben Tan, Pengyue Zhu, Jianwei Yang, Zhengyuan Yang, Lijuan Wang, Andrey Kolobov, Jianfeng Gao, Boqing Gong

机构 * Boston University(波士顿大学) Microsoft Research(微软研究院)

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05469 2025-11-14 cs.CV

Generating Physically Stable and Buildable Brick Structures from Text

Ava Pun, Kangle Deng, Ruixuan Liu, Deva Ramanan, Changliu Liu, Jun-Yan Zhu

机构 * Carnegie Mellon University(卡内基梅隆大学)

Comments Project page: https://avalovelace1.github.io/BrickGPT/

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 14798-14809

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.13395 2025-11-13 cs.CV

ACDC: The Adverse Conditions Dataset with Correspondences for Robust Semantic Driving Scene Perception

Christos Sakaridis, Haoran Wang, Ke Li, René Zurbrügg, Arpit Jadon, Wim Abbeloos, Daniel Olmeda Reino, Luc Van Gool, Dengxin Dai

机构 * ETH Zürich(苏黎世联邦理工学院) Max Planck Institute for Informatics(马克斯·普朗克信息研究所) Toyota Motor Europe(丰田欧洲公司) INSAIT Sofia University St. Kliment Ohridski(索菲亚大学圣克莱门特·奥赫里迪斯)

Comments IEEE T-PAMI 2025. Extended version of original conference paper published in ICCV 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16975 2025-11-13 cs.CV

EgoTV: Egocentric Task Verification from Natural Language Task Descriptions

Rishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra, Ruta Desai

机构 * Örebro University(奥雷布罗大学) Meta

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07409 2025-11-11 cs.CV

DIMO: Diverse 3D Motion Generation for Arbitrary Objects

Linzhan Mou, Jiahui Lei, Chen Wang, Lingjie Liu, Kostas Daniilidis

机构 * University of Pennsylvania(宾夕法尼亚大学) Archimedes, Athena RC(Archimedes,Athena RC)

Comments Published in ICCV 2025, project page https://linzhanm.github.io/dimo

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04781 2025-11-11 cs.LG

FedPall: Prototype-based Adversarial and Collaborative Learning for Federated Learning with Feature Drift

Yong Zhang, Feng Liang, Guanghu Yuan, Min Yang, Chengming Li, Xiping Hu

机构 * Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(深圳MSU-BIT大学人工智能研究院) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) School of Medical Technology, Beijing Institute of Technology(北京理工大学医学技术学院)

Comments 10 pages, 6 figures, and 1 table

Journal ref International Conference on Computer Vision (ICCV), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05664 2025-11-11 cs.CV cs.LG

Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model Unlearning

Saemi Moon, Minjong Lee, Sangdon Park, Dongwoo Kim

机构 * CSE, POSTECH(POSTECH 计算科学与工程系) GSAI, POSTECH(POSTECH 人工智能研究所)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06272 2025-11-11 cs.CV cs.AI

LaneDiffusion: Improving Centerline Graph Learning via Prior Injected BEV Feature Generation

Zijie Wang, Weiming Zhang, Wei Zhang, Xiao Tan, Hongxing Liu, Yaowei Wang, Guanbin Li

机构 * Sun Yat-sen University(中山大学) Shenzhen Loop Area Institute(深圳河套学院) Baidu Inc.(百度公司) Harbin Institute of Technology(哈尔滨工业大学) Pengcheng Laboratory(鹏城实验室) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06016 2025-11-11 cs.CV cs.AI

One-Shot Knowledge Transfer for Scalable Person Re-Identification

Longhua Li, Lei Qi, Xin Geng

机构 * School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学,南京,中国) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用关键实验室(东南大学),教育部,中国)

Comments Accepted by ICCV 2025

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04090 2025-11-11 cs.CV

Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework

Yi-Ting Chen, Ting-Hsuan Liao, Pengsheng Guo, Alexander Schwing, Jia-Bin Huang

机构 * University of Maryland, College Park(马里兰大学 College Park 分校) Carnegie Mellon University(卡内基梅隆大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

Comments Accepted to ICCV 2025. Project website: https://consistent3dsr.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19604 2025-11-11 eess.IV cs.CV

X-Diffusion: Generating Detailed 3D MRI Volumes From a Single Image Using Cross-Sectional Diffusion Models

Emmanuelle Bourigault, Abdullah Hamdi, Amir Jamaludin

机构 * Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

Comments accepted at ICCV 2025 GAIA workshop https://era-ai-biomed.github.io/GAIA/ , project website: https://emmanuelleb985.github.io/XDiffusion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20349 2025-11-10 cs.CV

Consistency Trajectory Matching for One-Step Generative Super-Resolution

Weiyi You, Mingyang Zhang, Leheng Zhang, Xingyu Zhou, Kexuan Shi, Shuhang Gu

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14516 2025-11-07 cs.CV

Back on Track: Bundle Adjustment for Dynamic Scene Reconstruction

Weirong Chen, Ganlin Zhang, Felix Wimbauer, Rui Wang, Nikita Araslanov, Andrea Vedaldi, Daniel Cremers

机构 * TU Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) University of Oxford(牛津大学) Microsoft(微软公司)

Comments ICCV 2025 Oral. Project page: https://wrchen530.github.io/projects/batrack/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18602 2025-11-07 eess.IV cs.CV

Evaluating and Improving the Effectiveness of Synthetic Chest X-Rays for Medical Image Analysis

Eva Prakash, Jeya Maria Jose Valanarasu, Zhihong Chen, Eduardo Pontes Reis, Andrew Johnston, Anuj Pareek, Christian Bluethgen, Sergios Gatidis, Cameron Olsen, Akshay Chaudhari, Andrew Ng, Curtis Langlotz

机构 * Stanford University(斯坦福大学)

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, October 2025, pages 4413-4421

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17341 2025-11-06 cs.LG cs.CR cs.CY cs.DC cs.ET

MetaFed: Advancing Privacy, Performance, and Sustainability in Federated Metaverse Systems

Muhammet Anil Yagiz, Zeynep Sude Cengiz, Polat Goktas

机构 * Department of Computer Engineering(计算机工程系) Faculty of Engineering(工程学院) Kırıkkale University(Kirikkkale 大学) Natural Sciences, Sabancı University(自然科学,萨班奇大学)

Comments 2025 IEEE International Symposium on Emerging Metaverse (ISEMV), co-located with the 2025 IEEE/CVF International Conference on Computer Vision (ICCV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18422 2025-11-06 cs.CV

Breaking the Encoder Barrier for Seamless Video-Language Understanding

Handong Li, Yiyuan Zhang, Longteng Guo, Xiangyu Yue, Jing Liu

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MMLab, CUHK(CUHK MMLab) Institute of Automation, Chinese Academy of Science(中国科学院自动化研究所) Shanghai AI Lab(上海人工智能实验室)

Comments 12 pages

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 23167-23176

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00011 2025-11-06 cs.LG cs.CV

DeepVAT: A Self-Supervised Technique for Cluster Assessment in Image Datasets

Alokendu Mazumder, Tirthajit Baruah, Akash Kumar Singh, Pagadla Krishna Murthy, Vishwajeet Pattanaik, Punit Rathore

Comments Accepted at ViPriors @ ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07661 2025-11-06 cs.CV cs.RO

ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones

Anurag Ghosh, Shen Zheng, Robert Tamburo, Khiem Vuong, Juan Alvarez-Padilla, Hailiang Zhu, Michael Cardei, Nicholas Dunn, Christoph Mertz, Srinivasa G. Narasimhan

机构 * Carnegie Mellon University(卡内基梅隆大学)

Comments ICCV 2025 Accepted Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01999 2025-11-05 cs.RO cs.AI

TRACE: Textual Reasoning for Affordance Coordinate Extraction

Sangyun Park, Jin Kim, Yuchen Cui, Matthew S. Brown

机构 * ABB Robotics(ABB机器人公司) University of California, Los Angeles(加州大学洛杉矶分校)

Comments ICCV 2025. *Equal contribution. †Corresponding author

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01175 2025-11-05 cs.CV

Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-Resolution

Peng Du, Hui Li, Han Xu, Paul Barom Jeon, Dongwook Lee, Daehyun Ji, Ran Yang, Feng Zhu

机构 * Samsung R&D Institute China Xi’an (SRCX)(三星中国研发中心西安(SRCX)) Samsung Electronics Co., LTD., South Korea(三星电子有限公司,韩国)

Comments ICCV 2025 Oral Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07215 2025-11-05 cs.RO cs.MM

RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation

Feng Yan, Fanfan Liu, Liming Zheng, Yufeng Zhong, Yiyang Huang, Zechao Guan, Chengjian Feng, Lin Ma

机构 * Meituan(美团)

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00502 2025-11-04 cs.CV cs.CL

ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers

Qianhao Yuan, Qingyu Zhang, Yanjiang Liu, Jiawei Chen, Yaojie Lu, Hongyu Lin, Jia Zheng, Xianpei Han, Le Sun

机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学)

Comments Published as a conference paper at ICCV 2025. Project page: https://github.com/icip-cas/ShortV

详情

展开后加载摘要…

URL PDF HTML 收藏