arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作者

Trevor Darrell

Computer Vision

至 收录 357
2504.13152 2025-04-18 cs.CV

St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World

Haiwen Feng, Junyi Zhang, Qianqian Wang, Yufei Ye, Pengcheng Yu, Michael J. Black, Trevor Darrell, Angjoo Kanazawa

Comments Project page: https://St4RTrack.github.io/

URL PDF HTML 收藏
2209.08763 2025-04-17 cs.RO cs.CV

Decentralized Vehicle Coordination: The Berkeley DeepDrive Drone Dataset and Consensus-Based Models

Fangyu Wu, Dequan Wang, Minjune Hwang, Chenhui Hao, Jiawei Lu, Jiamu Zhang, Christopher Chou, Trevor Darrell, Alexandre Bayen

Comments 7 pages, 7 figures, 1 table

URL PDF HTML 收藏
2412.03572 2025-04-15 cs.CV cs.AI cs.LG cs.RO

Navigation World Models

Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, Yann LeCun

Comments CVPR 2025. Project page: https://www.amirbar.net/nwm/

URL PDF HTML 收藏
2401.14391 2025-04-11 cs.CV

Rethinking Patch Dependence for Masked Autoencoders

Letian Fu, Long Lian, Renhao Wang, Baifeng Shi, Xudong Wang, Adam Yala, Trevor Darrell, Alexei A. Efros, Ken Goldberg

Comments Transactions on Machine Learning Research (TMLR) 2025

URL PDF HTML 收藏
2503.15485 2025-04-09 cs.CV cs.AI cs.CL cs.LG

TULIP: Towards Unified Language-Image Pretraining

Zineng Tang, Long Lian, Seun Eisape, XuDong Wang, Roei Herzig, Adam Yala, Alane Suhr, Trevor Darrell, David M. Chan

Comments (v2) Clarified fine-tuning process, updated appendix

URL PDF HTML 收藏
2412.08687 2025-03-27 cs.CV

VisionArena: 230K Real World User-VLM Conversations with Preference Labels

Christopher Chou, Lisa Dunlap, Koki Mashita, Krishna Mandal, Trevor Darrell, Ion Stoica, Joseph E. Gonzalez, Wei-Lin Chiang

Comments updated for CVPR Camera Ready

URL PDF HTML 收藏
2407.18908 2025-03-21 cs.LG cs.CL cs.CV

Wolf: Dense Video Captioning with a World Summarization Framework

Boyi Li, Ligeng Zhu, Ran Tian, Shuhan Tan, Yuxiao Chen, Yao Lu, Yin Cui, Sushant Veer, Max Ehrlich, Jonah Philion, Xinshuo Weng, Fuzhao Xue, Linxi Fan, Yuke Zhu, Jan Kautz, Andrew Tao, Ming-Yu Liu, Sanja Fidler, Boris Ivanovic, Trevor Darrell, Jitendra Malik, Song Han, Marco Pavone

URL PDF HTML 收藏
2503.12355 2025-03-18 cs.CV cs.LG

Atlas: Multi-Scale Attention Improves Long Context Image Modeling

Kumar Krishna Agrawal, Long Lian, Longchao Liu, Natalia Harguindeguy, Boyi Li, Alexander Bick, Maggie Chung, Trevor Darrell, Adam Yala

URL PDF HTML 收藏
2410.12782 2025-03-18 cs.RO cs.CL

In-Context Learning Enables Robot Action Prediction in LLMs

Yida Yin, Zekai Wang, Yuvan Sharma, Dantong Niu, Trevor Darrell, Roei Herzig

Comments Published in ICRA 2025

URL PDF HTML 收藏
2311.05589 2025-03-18 cs.LG math.OC stat.ML

A Coefficient Makes SVRG Effective

Yida Yin, Zhiqiu Xu, Zhiyuan Li, Trevor Darrell, Zhuang Liu

Comments Published in ICLR 2025

URL PDF HTML 收藏
2410.19314 2025-03-14 cs.CY cs.CL

Revealing and Reducing Gender Biases in Vision and Language Assistants (VLAs)

Leander Girrbach, Stephan Alaniz, Yiran Huang, Trevor Darrell, Zeynep Akata

Comments Accepted at ICLR 2025

URL PDF HTML 收藏
2503.07860 2025-03-12 cs.CV cs.AI cs.LG

Video Action Differencing

James Burgess, Xiaohan Wang, Yuhui Zhang, Anita Rau, Alejandro Lozano, Lisa Dunlap, Trevor Darrell, Serena Yeung-Levy

Comments ICLR 2025 (International Conference on Learning Representations) Project page: http://jmhb0.github.io/viddiff Benchmark: https://huggingface.co/datasets/jmhb/VidDiffBench

URL PDF HTML 收藏
2407.13766 2025-03-12 cs.CV

Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Tsung-Han Wu, Giscard Biamby, Jerome Quenum, Ritwik Gupta, Joseph E. Gonzalez, Trevor Darrell, David M. Chan

Comments Accepted to ICLR 2025; Project page: https://visual-haystacks.github.io

URL PDF HTML 收藏
2503.06469 2025-03-11 cs.CV

Vector Quantized Feature Fields for Fast 3D Semantic Lifting

George Tang, Aditya Agarwal, Weiqiao Han, Trevor Darrell, Yutong Bai

URL PDF HTML 收藏
2402.13144 2025-01-03 cs.LG cs.CV

Neural Network Diffusion

Kai Wang, Dongwen Tang, Boya Zeng, Yida Yin, Zhaopan Xu, Yukun Zhou, Zelin Zang, Trevor Darrell, Zhuang Liu, Yang You

Comments We introduce a novel approach for parameter generation, named neural network parameter diffusion (\textbf{p-diff}), which employs a standard latent diffusion model to synthesize a new set of parameters

URL PDF HTML 收藏
2406.15334 2024-12-23 cs.CV cs.AI cs.CL cs.LG

Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning

Brandon Huang, Chancharik Mitra, Assaf Arbelle, Leonid Karlinsky, Trevor Darrell, Roei Herzig

Comments Published in NeurIPS 2024

URL PDF HTML 收藏
2412.06774 2024-12-10 cs.CV cs.AI cs.LG

Visual Lexicon: Rich Image Features in Language Space

XuDong Wang, Xingyi Zhou, Alireza Fathi, Trevor Darrell, Cordelia Schmid

Comments Tech report. 16 pages, 10 figures

URL PDF HTML 收藏
2406.08164 2024-11-14 cs.CV

ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs

Irene Huang, Wei Lin, M. Jehanzeb Mirza, Jacob A. Hansen, Sivan Doveh, Victor Ion Butoi, Roei Herzig, Assaf Arbelle, Hilde Kuehne, Trevor Darrell, Chuang Gan, Aude Oliva, Rogerio Feris, Leonid Karlinsky

Comments NeurIPS 2024 Camera Ready

URL PDF HTML 收藏
2411.05001 2024-11-08 cs.CV cs.AI cs.CL cs.LG

Analyzing The Language of Visual Tokens

David M. Chan, Rodolfo Corona, Joonyong Park, Cheol Jun Cho, Yutong Bai, Trevor Darrell

URL PDF HTML 收藏
2410.18923 2024-11-04 cs.CV cs.AI

SegLLM: Multi-round Reasoning Segmentation

XuDong Wang, Shaolun Zhang, Shufan Li, Konstantinos Kallidromitis, Kehan Li, Yusuke Kato, Kazuki Kozuka, Trevor Darrell

Comments 22 pages, 10 figures, 11 tables

URL PDF HTML 收藏
2404.01476 2024-10-22 cs.CV cs.AI cs.CL cs.LG

TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering

Chuyi Shang, Amos You, Sanjay Subramanian, Trevor Darrell, Roei Herzig

Comments EMNLP 2024 (Main)

URL PDF HTML 收藏
2410.10817 2024-10-15 cs.CV cs.LG

When Does Perceptual Alignment Benefit Vision Representations?

Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Netanel Y. Tamir, Lucy Chai, Simon Kornblith, Trevor Darrell, Phillip Isola

Comments S.S. and S.F. contributed equally. Website: percep-align.github.io

URL PDF HTML 收藏
2404.05729 2024-10-08 cs.CV

Finding Visual Task Vectors

Alberto Hojel, Yutong Bai, Trevor Darrell, Amir Globerson, Amir Bar

Comments https://github.com/alhojel/visual_task_vectors

URL PDF HTML 收藏
2410.03654 2024-10-07 cs.RO cs.LG

Learning Humanoid Locomotion over Challenging Terrain

Ilija Radosavovic, Sarthak Kamat, Trevor Darrell, Jitendra Malik

Comments Project page: https://humanoid-challenging-terrain.github.io

URL PDF HTML 收藏
2409.17216 2024-09-27 cs.CY cs.AI

Data-Centric AI Governance: Addressing the Limitations of Model-Focused Policies

Ritwik Gupta, Leah Walker, Rodolfo Corona, Stephanie Fu, Suzanne Petryk, Janet Napolitano, Trevor Darrell, Andrew W. Reddie

URL PDF HTML 收藏
2403.01915 2024-07-23 cs.CV cs.AI

xT: Nested Tokenization for Larger Context in Large Images

Ritwik Gupta, Shufan Li, Tyler Zhu, Jitendra Malik, Trevor Darrell, Karttikeya Mangalam

Comments Accepted to the 2024 International Conference on Machine Learning (ICML)

URL PDF HTML 收藏
2403.13043 2024-07-19 cs.CV

When Do We Not Need Larger Vision Models?

Baifeng Shi, Ziyang Wu, Maolin Mao, Xin Wang, Trevor Darrell

Comments Code: https://github.com/bfshi/scaling_on_scales

URL PDF HTML 收藏
2312.02249 2024-07-11 cs.CV cs.CL

Recursive Visual Programming

Jiaxin Ge, Sanjay Subramanian, Baifeng Shi, Roei Herzig, Trevor Darrell

URL PDF HTML 收藏
2406.20081 2024-07-01 cs.CV cs.LG

Segment Anything without Supervision

XuDong Wang, Jingfeng Yang, Trevor Darrell

Comments Code: https://github.com/frank-xwang/UnSAM

URL PDF HTML 收藏
2406.11815 2024-06-18 cs.RO cs.CV cs.LG

LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Dantong Niu, Yuvan Sharma, Giscard Biamby, Jerome Quenum, Yutong Bai, Baifeng Shi, Trevor Darrell, Roei Herzig

URL PDF HTML 收藏