I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
John Burden, Jonathan Prunty, Ben Slater, Matthieu Tehenan, Greg Davis, Lucy Cheke
机构
*
Leverhulme Centre for the Future of Intelligence, University of Cambridge(未来智能研究中心、剑桥大学)
;
Department of Engineering, University of Cambridge(工程系、剑桥大学)
;
Department of Psychology, University of Cambridge(心理学系、剑桥大学)
;
Department of Computer Science, University of Cambridge(计算机科学系、剑桥大学)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.CV、cs.AI
GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
Qifu Wen, Xi Zeng, Zihan Zhou, Shuaijun Liu, Mehdi Hosseinzadeh, Ningxin Su, Reza Rawassizadeh
机构
*
Department of Computer Science, Boston University Metropolitan College(波士顿大学计算机科学系)
;
Information Hub, The Hong Kong University of Science and Technology, Guangzhou(香港科技大学广州信息中心)
;
School of Engineering and Technology, Duy Tan University, Da Nang, Vietnam(杜益大学工程科技学院,岘港,越南)
;
Department of AI, School of Computer Science and Engineering, Galgotias University, Greater Noida, India(加洛吉亚大学人工智能系,诺伊达,印度)
Zekun Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jiashuo Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang
机构
*
Beihang University(北航)
;
M-A-P
;
The Hong Kong Polytechnic University(香港理工大学)
;
AIWaves
;
University of Alberta(阿尔伯塔大学)
;
University of Waterloo(滑铁卢大学)
;
University of Manchester(曼彻斯特大学)
;
Chinese Academy of Sciences(中国科学院)
;
Peking University(北京大学)
;
Shanghai AI Lab(上海AI实验室)
;
Nanjing University(南京大学)
;
Kuaishou Technology(快手科技)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.AI、cs.LG
机构
*
Department of Electrical and Electronic Engineering, The University of Hong Kong(香港大学电子与电气工程系)
;
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
;
NVIDIA Research(NVIDIA研究)
Stacked Regression using Off-the-shelf, Stimulus-tuned and Fine-tuned Neural Networks for Predicting fMRI Brain Responses to Movies (Algonauts 2025 Report)
Robert Scholz, Kunal Bagga, Christine Ahrends, Carlo Alberto Barbano
机构
*
Université Paris Cité(巴黎Cité大学)
;
University of Oxford(牛津大学)
;
University of Turin(都灵大学)
;
Universität Leipzig(莱比锡大学)
;
Max Planck School of Cognition(马克斯·普朗克认知科学学院)
Yang Chen, Yanbin Wei, Ke Jin, Yi Kong, James Kwok, Yu Zhang
机构
*
Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系)
;
Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)
;
School of Information and Control Engineering, China University of Mining and Technology(中国矿业大学信息与控制工程学院)
Random Direct Preference Optimization for Radiography Report Generation
Valentin Samokhin, Boris Shirokikh, Mikhail Goncharov, Dmitriy Umerenkov, Maksim Bobrin, Ivan Oseledets, Dmitry Dylov, Mikhail Belyaev
机构
*
Artificial Intelligence Research Institute (AIRI)(人工智能研究院)
;
Skolkovo Institute of Science and Technology(斯克罗夫学院)
;
Institute for Information Transmission Problems(信息传输问题研究所)
专题命中
VLM训练与架构
:visual language model(abstract);分类 cs.CV、cs.AI
Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
Remo Sasso, Michelangelo Conserva, Dominik Jeurissen, Paulo Rauber
机构
*
School of Electronic Engineering and Computer Science(电子工程与计算机科学学院)
专题命中
VLM训练与架构
:VLM(abstract);分类 cs.AI、cs.LG
Comments16 pages, 7 figures. Accepted for presentation at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on the Foundations of Reasoning in Language Models (FoRLM)