CommentsThis paper is being withdrawn because we have identified a significant error in the implementation of our self-supervised clustering approach. Specifically, our feature aggregation step inadvertently leaked temporal information across frames, which violates the core assumption of our training-free method. We sincerely apologize to the research community
Caption This, Reason That: VLMs Caught in the Middle
Zihan Weng, Lucas Gomez, Taylor Whittington Webb, Pouya Bashivan
机构
*
Integrated Program in Neuroscience (IPN) McGill University(神经科学联合计划 麦吉尔大学)
;
Mila, University of Montreal(蒙特利尔大学Mila)
;
Microsoft Research USA(微软研究院美国总部)
;
Department of Physiology McGill University(生理学系 麦吉尔大学)
Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
Sahar Salimpour, Lei Fu, Kajetan Rachwał, Pascal Bertrand, Kevin O'Sullivan, Robert Jakob, Farhad Keramat, Leonardo Militano, Giovanni Toffetti, Harry Edelman, Jorge Peña Queralta
机构
*
Department of Computing, University of Turku(图尔库大学计算机系)
;
Institute of Computer Science, Zurich University of Applied Sciences(应用科学大学计算机科学研究所)
;
Centre for Artificial Ingelligence, Zurich University of Applied Sciences(应用科学大学人工智能中心)
;
Agentic Systems Lab, Department of Management, Technology and Economics, ETH Zürich(苏黎世联邦理工学院管理、科技与经济系代理系统实验室)
;
Faculty of Mathematics and Information Science, Warsaw University of Technology(华沙技术大学数学与信息科学学院)
专题命中
评测与基准
:LLM(title);large language model(abstract);language model(abstract);foundation model(abstract)
Captions Speak Louder than Images: Generalizing Foundation Models for E-commerce from High-quality Multimodal Instruction Data
Xinyi Ling, Hanwen Du, Bo Peng, Zhihui Zhu, Xia Ning
机构
*
Department of Computer Science and Engineering, The Ohio State University(计算机科学与工程系,俄亥俄州立大学)
;
Translational Data Analytics Institute, The Ohio State University(转化数据分析研究所,俄亥俄州立大学)
;
Department of Biomedical Informatics, The Ohio State University(生物医学信息学系,俄亥俄州立大学)
机构
*
Indian Institute of Technology Mandi(印度理工学院曼迪分校)
;
Vellore Institute of Technology(韦洛雷理工学院)
;
Indian Institute of Technology Kharagpur(印度理工学院哈里科普分校)
PALMS+: Modular Image-Based Floor Plan Localization Leveraging Depth Foundation Model
Yunqian Cheng, Benjamin Princen, Roberto Manduchi
机构
*
University of California, Santa Cruz(加州大学圣克ruz分校)
专题命中
评测与基准
:foundation model(title);分类 cs.AI
CommentsAccepted to IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026, Application Track. Main paper: 8 pages, 5 figures. Supplementary material included
Querying Labeled Time Series Data with Scenario Programs
Edward Kim, Devan Shanker, Varun Bharadwaj, Hongbeen Park, Jinkyu Kim, Hazem Torfah, Daniel J Fremont, Sanjit A Seshia
机构
*
University of California, Berkeley(加州大学伯克利分校)
;
Korea University(韩国大学)
;
Chalmers University of Technology(查尔姆斯理工大学)
;
University of Gothenburg(哥德堡大学)
;
University of California, Santa Cruz(加州大学圣克ruz分校)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
MMTEB: Massive Multilingual Text Embedding Benchmark
Kenneth Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzemiński, Genta Indra Winata, Saba Sturua, Saiteja Utpala, Mathieu Ciancone, Marion Schaeffer, Gabriel Sequeira, Diganta Misra, Shreeya Dhakal, Jonathan Rystrøm, Roman Solomatin, Ömer Çağatan, Akash Kundu, Martin Bernstorff, Shitao Xiao, Akshita Sukhlecha, Bhavish Pahwa, Rafał Poświata, Kranthi Kiran GV, Shawon Ashraf, Daniel Auras, Björn Plüster, Jan Philipp Harries, Loïc Magne, Isabelle Mohr, Mariya Hendriksen, Dawei Zhu, Hippolyte Gisserot-Boukhlef, Tom Aarsen, Jan Kostkan, Konrad Wojtasik, Taemin Lee, Marek Šuppa, Crystina Zhang, Roberta Rocca, Mohammed Hamdy, Andrianos Michail, John Yang, Manuel Faysse, Aleksei Vatolin, Nandan Thakur, Manan Dey, Dipam Vasani, Pranjal Chitale, Simone Tedeschi, Nguyen Tai, Artem Snegirev, Michael Günther, Mengzhou Xia, Weijia Shi, Xing Han Lù, Jordan Clive, Gayatri Krishnakumar, Anna Maksimova, Silvan Wehrli, Maria Tikhonova, Henil Panchal, Aleksandr Abramov, Malte Ostendorff, Zheng Liu, Simon Clematide, Lester James Miranda, Alena Fenogenova, Guangyu Song, Ruqiya Bin Safi, Wen-Ding Li, Alessia Borghini, Federico Cassano, Hongjin Su, Jimmy Lin, Howard Yen, Lasse Hansen, Sara Hooker, Chenghao Xiao, Vaibhav Adlakha, Orion Weller, Siva Reddy, Niklas Muennighoff
机构
*
Aarhus University(奥胡斯大学)
;
Individual Contributor(个人贡献者)
;
Esker(Esker公司)
;
INSA Lyon(里昂INSA)
;
University of Amsterdam(阿姆斯特丹大学)
;
MBZUAI(穆罕默德·本·拉希德智能技术研究院)
;
Jina AI(Jina AI公司)
;
Microsoft Research(微软研究院)
;
Wikit(Wikit公司)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
Mohsinul Kabir, Ajwad Abrar, Sophia Ananiadou
机构
*
Department of Computer Science, National Center for Text Mining, The University of Manchester(计算机科学系,文本挖掘国家中心,曼彻斯特大学)
;
Department of Computer Science and Engineering, Islamic University of Technology(计算机科学与工程系,伊斯兰技术大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
CommentsAccepted at EMNLP 2025 (Main)
Journal refProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
EcomMMMU: Strategic Utilization of Visuals for Robust Multimodal E-commerce Models
Xinyi Ling, Hanwen Du, Zhihui Zhu, Xia Ning
机构
*
Department of Computer Science and Engineering, The Ohio State University(俄亥俄州立大学计算机科学与工程系)
;
Translational Data Analytics Institute, The Ohio State University(俄亥俄州立大学转化数据分析研究所)
;
Department of Biomedical Informatics, The Ohio State University(俄亥俄州立大学生物医学信息学系)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI