arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作者

Cordelia Schmid

Computer Vision

至 收录 226
2504.06021 2025-04-09 cs.CV

Memory-Modular Classification: Learning to Generalize with Memory Replacement

Dahyun Kang, Ahmet Iscen, Eunchan Jo, Sua Choi, Minsu Cho, Cordelia Schmid

Comments Accepted to TMLR. Code available: https://github.com/dahyun-kang/mml

URL PDF HTML 收藏
2504.05303 2025-04-08 cs.CV

InteractVLM: 3D Interaction Reasoning from 2D Foundational Models

Sai Kumar Dwivedi, Dimitrije Antić, Shashank Tripathi, Omid Taheri, Cordelia Schmid, Michael J. Black, Dimitrios Tzionas

Comments CVPR 2025

URL PDF HTML 收藏
2412.05796 2025-04-08 cs.CV cs.AI cs.LG

Language-Guided Image Tokenization for Generation

Kaiwen Zha, Lijun Yu, Alireza Fathi, David A. Ross, Cordelia Schmid, Dina Katabi, Xiuye Gu

Comments CVPR 2025 Oral. Project page: https://kaiwenzha.github.io/textok/

URL PDF HTML 收藏
2308.14746 2025-04-02 cs.CV

CoVR-2: Automatic Data Construction for Composed Video Retrieval

Lucas Ventura, Antoine Yang, Cordelia Schmid, Gül Varol

Comments Appears in TPAMI 2024 (DOI: 10.1109/TPAMI.2024.3463799). Journal extension of the AAAI 2024 conference paper arXiv:2308.14746v3. Project page: https://imagine.enpc.fr/~ventural/covr/

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

URL PDF HTML 收藏
2504.00072 2025-04-02 cs.CV

Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs

Lucas Ventura, Antoine Yang, Cordelia Schmid, Gül Varol

Comments CVPR 2025 Camera ready. Project page: https://imagine.enpc.fr/~lucas.ventura/chapter-llama/

URL PDF HTML 收藏
2404.06511 2025-03-28 cs.CV cs.AI cs.LG

MoReVQA: Exploring Modular Reasoning Models for Video Question Answering

Juhong Min, Shyamal Buch, Arsha Nagrani, Minsu Cho, Cordelia Schmid

Comments CVPR 2024; updated NExT-GQA results in Appendix

URL PDF HTML 收藏
2503.18897 2025-03-25 cs.CV cs.RO

Online 3D Scene Reconstruction Using Neural Object Priors

Thomas Chabal, Shizhe Chen, Jean Ponce, Cordelia Schmid

Comments 3DV 2025. Project page: https://www.di.ens.fr/willow/research/online-scene-reconstruction/

URL PDF HTML 收藏
2407.13579 2025-03-12 cs.CL

Towards Zero-Shot Multimodal Machine Translation

Matthieu Futeral, Cordelia Schmid, Benoît Sagot, Rachel Bawden

Comments NAACL 2025 (Findings)

URL PDF HTML 收藏
2503.04919 2025-03-10 cs.CV

FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement

Ian Huang, Yanan Bao, Karen Truong, Howard Zhou, Cordelia Schmid, Leonidas Guibas, Alireza Fathi

URL PDF HTML 收藏
2503.04666 2025-03-07 cs.CV

What Are You Doing? A Closer Look at Controllable Human Video Generation

Emanuele Bugliarello, Anurag Arnab, Roni Paiss, Pieter-Jan Kindermans, Cordelia Schmid

URL PDF HTML 收藏
2410.01345 2025-03-04 cs.RO cs.CV

Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy

Ricardo Garcia, Shizhe Chen, Cordelia Schmid

Comments ICRA 2025

URL PDF HTML 收藏
2404.15709 2025-03-04 cs.CV cs.LG cs.RO

ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos

Zerui Chen, Shizhe Chen, Etienne Arlaud, Ivan Laptev, Cordelia Schmid

Comments Accepted by ICRA 2025. Project Page: https://zerchen.github.io/projects/vividex.html

URL PDF HTML 收藏
2412.09582 2025-01-22 cs.LG cs.AI cs.CV

Neptune: The Long Orbit to Benchmarking Long Video Understanding

Arsha Nagrani, Mingda Zhang, Ramin Mehran, Rachel Hornung, Nitesh Bharadwaj Gundavarapu, Nilpa Jha, Austin Myers, Xingyi Zhou, Boqing Gong, Cordelia Schmid, Mikhail Sirotenko, Yukun Zhu, Tobias Weyand

URL PDF HTML 收藏
2412.06774 2024-12-10 cs.CV cs.AI cs.LG

Visual Lexicon: Rich Image Features in Language Space

XuDong Wang, Xingyi Zhou, Alireza Fathi, Trevor Darrell, Cordelia Schmid

Comments Tech report. 16 pages, 10 figures

URL PDF HTML 收藏
2411.07584 2024-11-13 cs.CV

Grounded Video Caption Generation

Evangelos Kazakos, Cordelia Schmid, Josef Sivic

URL PDF HTML 收藏
2410.23676 2024-11-01 cs.CV

Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach

Mathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet Iscen

Comments NeurIPS 2024

URL PDF HTML 收藏
2306.11729 2024-10-16 cs.CV

Dense Video Object Captioning from Disjoint Supervision

Xingyi Zhou, Anurag Arnab, Chen Sun, Cordelia Schmid

Comments Code is available at https://github.com/google-research/scenic/tree/main/scenic/projects/densevoc

URL PDF HTML 收藏
2401.06035 2024-08-13 cs.CV cs.LG

RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks

Partha Ghosh, Soubhik Sanyal, Cordelia Schmid, Bernhard Schölkopf

URL PDF HTML 收藏
2304.06372 2024-07-23 cs.RO

Contact Models in Robotics: a Comparative Analysis

Quentin Le Lidec, Wilson Jallet, Louis Montaut, Ivan Laptev, Cordelia Schmid, Justin Carpentier

URL PDF HTML 收藏
2407.10910 2024-07-17 cs.CV cs.LG

DataDream: Few-shot Guided Dataset Generation

Jae Myung Kim, Jessica Bader, Stephan Alaniz, Cordelia Schmid, Zeynep Akata

Comments Accepted to ECCV 2024

URL PDF HTML 收藏
2404.17498 2024-04-29 cs.CV

Learning text-to-video retrieval from image captioning

Lucas Ventura, Cordelia Schmid, Gül Varol

Comments A short version of this work appeared at CVPR 2023 Workshops. Project page: https://imagine.enpc.fr/~ventural/multicaps/

URL PDF HTML 收藏
2404.03924 2024-04-08 cs.CV

Learning Correlation Structures for Vision Transformers

Manjin Kim, Paul Hongsuck Seo, Cordelia Schmid, Minsu Cho

Comments Accepted to CVPR 2024

URL PDF HTML 收藏
2404.01491 2024-04-03 cs.CV

SUGAR: Pre-training 3D Visual Representations for Robotics

Shizhe Chen, Ricardo Garcia, Ivan Laptev, Cordelia Schmid

Comments Accepted to CVPR 2024. Project webpage: https://cshizhe.github.io/projects/robot_sugar.html

URL PDF HTML 收藏
2404.01297 2024-04-02 cs.CV

Streaming Dense Video Captioning

Xingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan, Austin Myers, Xuehan Xiong, Arsha Nagrani, Cordelia Schmid

Comments CVPR 2024. Code is available at https://github.com/google-research/scenic/tree/main/scenic/projects/streaming_dvc

URL PDF HTML 收藏
2403.02041 2024-03-22 cs.CV

A Generative Approach for Wikipedia-Scale Visual Entity Recognition

Mathilde Caron, Ahmet Iscen, Alireza Fathi, Cordelia Schmid

Comments CVPR2024

URL PDF HTML 收藏
2312.00786 2024-03-05 cs.CV

Dense Optical Tracking: Connecting the Dots

Guillaume Le Moing, Jean Ponce, Cordelia Schmid

Comments Accepted to CVPR 2024

URL PDF HTML 收藏
2403.01248 2024-03-05 cs.CV cs.AI cs.CL cs.LG

SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code

Ziniu Hu, Ahmet Iscen, Aashi Jain, Thomas Kipf, Yisong Yue, David A. Ross, Cordelia Schmid, Alireza Fathi

URL PDF HTML 收藏
2306.07196 2024-02-22 cs.CV

Retrieval-Enhanced Contrastive Vision-Text Models

Ahmet Iscen, Mathilde Caron, Alireza Fathi, Cordelia Schmid

URL PDF HTML 收藏
2402.02887 2024-02-06 cs.CV cs.LG

Time-, Memory- and Parameter-Efficient Visual Adaptation

Otniel-Bogdan Mercea, Alexey Gritsenko, Cordelia Schmid, Anurag Arnab

URL PDF HTML 收藏
2203.03986 2024-01-23 cs.RO math.OC

Leveraging Randomized Smoothing for Optimal Control of Nonsmooth Dynamical Systems

Quentin Le Lidec, Fabian Schramm, Louis Montaut, Cordelia Schmid, Ivan Laptev, Justin Carpentier

URL PDF HTML 收藏