EAA: Automating materials characterization with vision language model agents
EAA: 用视觉语言模型代理自动化材料表征
Ming Du, Yanqi Luo, Srutarshi Banerjee, Michael Wojcik, Jelena Popovic, Mathew J. Cherukara
机构
*
Argonne National Laboratory(阿贡国家实验室)
;
Advanced Photon Source(先进光子源)
;
Data Science and Learning Division(数据科学与学习 division)
;
Department of Radiation Oncology(放射肿瘤学系)
;
Northwestern University(西北大学)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
在简单视觉规划任务中多模态大语言模型推理的分布外泛化
Yannic Neuhaus, Nicolas Flammarion, Matthias Hein, Francesco Croce
机构
*
Tübingen AI Center -- University of Tübingen(图宾根人工智能中心 -- 图宾根大学)
;
EPFL(瑞士联邦理工学院)
;
ELLIS Institute Finland -- Aalto University(芬兰ELLIS研究所 -- 阿尔托大学)
How Vision Becomes Language: A Layer-wise Information-Theoretic Analysis of Multimodal Reasoning
视觉如何转化为语言:多模态推理的逐层信息论分析
Hongxuan Wu, Yukun Zhang, Xueqing Zhou
机构
*
Duke kunshan University, Duke University(杜克昆山大学、杜克大学)
;
The Chinese University of Hong Kong, Hong Kong, China(香港中文大学)
;
Fudan University, Shanghai, China(复旦大学)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
EventMemAgent: 基于分层事件中心记忆的在线视频理解与自适应工具使用
Siwei Wen, Zhangcheng Wang, Xingjian Zhang, Lei Huang, Wenjun Wu
机构
*
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University(北京未来区块链与隐私计算先进创新中心,人工智能学院,北京航空航天大学)
;
Hangzhou International Innovation Institute, Beihang University(杭州国际创新研究院,北京航空航天大学)
;
Paradigm Inc.(4Paradigm公司)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
基于多模态大语言模型的零样本人-物交互检测
Shiyu Xuan, Dongkai Wang, Zechao Li, Jinhui Tang
机构
*
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
;
School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics(西南财经大学计算机与人工智能学院)
;
Nanjing Forestry University(南京林业大学)
Concept-Enhanced Multimodal RAG: Towards Interpretable and Accurate Radiology Report Generation
概念增强的多模态RAG:迈向可解释且准确的放射科报告生成
Marco Salmè, Federico Siciliano, Fabrizio Silvestri, Paolo Soda, Rosa Sicilia, Valerio Guarrasi
机构
*
Department of Engineering(工程系)
;
Research Unit of Artificial Intelligence and Computer Systems(人工智能与计算机系统研究单位)
;
Università Campus Bio-Medico of Roma(罗马大学生物医学校园)
;
Department of Computer, Control and Management Engineering(计算机、控制与管理工程系)
;
Sapienza University of Rome(罗马萨皮恩扎大学)
;
Department of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering(诊断与介入、辐射物理、生物医学工程系)
;
Umeå University(乌梅拉大学)
;
UniCamillus-Saint Camillus International University of Health Sciences(UniCamillus-圣卡米卢斯国际健康科学大学)
机构
*
School of Electronic Engineering, Xidian University(西安电子科技大学电子工程学院)
;
School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)
;
College of Computer and Information, Hohai University(河海大学计算机与信息学院)
;
Institute for Infocomm Research, A*STAR(A*STAR信息与通信研究所)
Can Multimodal LLMs Perform Time Series Anomaly Detection?
多模态大语言模型能否进行时间序列异常检测?
Xiongxiao Xu, Haoran Wang, Yueqing Liang, Philip S. Yu, Yue Zhao, Kai Shu
机构
*
Illinois Institute of Technology(伊利诺伊理工学院)
;
Emory University(埃默里大学)
;
University of Illinois Chicago(伊利诺伊大学香槟分校)
;
University of Southern California(南加州大学)
Sonal Kumar, Prem Seetharaman, Ke Chen, Oriol Nieto, Jiaqi Su, Zhepei Wang, Rithesh Kumar, Dinesh Manocha, Nicholas J. Bryan, Zeyu Jin, Justin Salamon
机构
*
University of Maryland, College Park, USA(美国马里兰大学 College Park 分校)
;
Adobe Research, USA(Adobe 研究院)
;
OpenAI, USA (work done while at Adobe)(OpenAI, USA)
Learning to Retrieve Navigable Candidates for Efficient Vision-and-Language Navigation
学习高效视觉-语言导航中的可导航候选检索
Shutian Gu, Chengkai Huang, Ruoyu Wang, Lina Yao
机构
*
University of New South Wales, Sydney, Australia(新南威尔士大学)
;
Macquarie University, Sydney, Australia(麦考瑞大学)
;
Data61, CSIRO, Sydney, Australia(Data61, CSIRO)
;
The School of Computer Science(计算机科学学院)
机构
*
Australian Institute for Machine Learning, Adelaide University(澳大利亚机器学习研究所,阿德莱德大学)
;
The University of Manchester(曼彻斯特大学)
;
Zhejiang University(浙江大学)
;
Agency for Science, Technology and Research (A*STAR)(科技研究局(A*STAR))
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract)
CARE Drive A Framework for Evaluating Reason-Responsiveness of Vision Language Models in Automated Driving
CARE Drive:一种用于评估视觉语言模型在自动驾驶中推理响应性的框架
Lucas Elbert Suryana, Farah Bierenga, Sanne van Buuren, Pepijn Kooij, Elsefien Tulleners, Federico Scari, Simeon Calvert, Bart van Arem, Arkady Zgonnikov
机构
*
organization= Department of Transport \& Planning, Faculty of Civil Engineering
;
Geosciences, Delft University of Technology
;
organization= Department of Cognitive Robotics, Faculty of Mechanical Engineering, Delft University of Technology
;
organization= Centre for Meaningful Human Control, Delft University of Technology
;
organization= Faculty of Mechanical Engineering, Delft University of Technology , city= Delft , country= The Netherlands
专题命中
幻觉与鲁棒性
:vision language model(title,abstract);分类 cs.CV、cs.AI
AI总结
CARE Drive提出了一种评估视觉语言模型在自动驾驶中推理响应性的框架,通过受控上下文变化比较基线和增强模型决策,验证人类原因对决策的影响。
Comments21 pages, on submission to Transportation Research Part C
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
面向视觉语言模型的层次化通用多模态攻击
Peng-Fei Zhang, Zi Huang
机构
*
School of Electrical Engineering and Computer Science, the University of Queensland(电气工程与计算机科学学院,昆士兰大学)
;
Department of Computer Science, City University of Hong Kong(计算机科学系,香港城市大学)
Annotation-Efficient Vision-Language Model Adaptation to the Polish Language Using the LLaVA Framework
利用LLaVA框架实现波兰语视觉-语言模型的高效标注
Grzegorz Statkiewicz, Alicja Dobrzeniecka, Karolina Seweryn, Aleksandra Krasnodębska, Karolina Piosek, Katarzyna Bogusz, Sebastian Cygert, Wojciech Kusa
LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge
LQA: 一种轻量级量化自适应框架,用于边缘设备上的视觉-语言模型
Xin Wang, Hong Jia, Hualin Zhou, Sheng Guang Wang, Yu Zhang, Ting Dang, Tao Gu
机构
*
School of Computing(计算机学院)
;
Information Systems, University of Melbourne, Melbourne, Australia(信息系统学院,墨尔本大学,澳大利亚墨尔本)
;
Faculty of Science, Auckland, New Zealand(科学学院,新西兰奥克兰)
;
School of Computing, University of Macquarie, Sydney, Australia(计算机学院,麦考瑞大学,澳大利亚悉尼)
GMAIL: Generative Modality Alignment for generated Image Learning
GMAIL: 生成模态对齐用于生成图像学习
Shentong Mo, Sukmin Yun
机构
*
Department of Machine Learning, CMU, USA(卡内基梅隆大学机器学习系)
;
Department of Machine Learning, MBZUAI, UAE(马斯克大学人工智能研究所)
;
Department of Artificial Intelligence, Hanyang University ERICA, South Korea(翰阳大学ERICA人工智能系)