R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
R4:在4D时空空间中为视觉语言模型引入检索增强推理
Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax
机构
*
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
Esslingen University of Applied Sciences(埃斯林根应用科学大学)
;
Dr. Ing. h.c. F. Porsche AG(德意志联邦汽车工业协会)
;
University of Michigan(密歇根大学)
;
Voxel51 Inc.(Voxel51公司)
A multimodal Bayesian Network for symptom-level depression and anxiety prediction from voice and speech data
一种多模态贝叶斯网络用于从语音和语音数据中预测症状层面的抑郁和焦虑
Agnes Norbury, George Fairs, Alexandra L. Georgescu, Matthew M. Nour, Emilia Molimpakis, Stefano Goria
机构
*
thymia Limited(thymia有限公司)
;
Institute of Psychiatry, Psychology & Neuroscience, King’s College London(心理学与神经科学研究院,伦敦国王学院)
;
Department of Psychiatry, University of Oxford(牛津大学精神病学系)
;
Max Planck UCL Centre for Computational Psychiatry and Ageing, University College London(Max Planck大学学院计算精神病学与衰老中心,伦敦大学学院)
CommentsThe method has flaws, especially with the decoupling module. During the decoupling process, the heterogeneity of the three modal data and the differences in distribution were not taken into account
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
CoVAR: 通过多模态扩散生成视频与动作用于机器人操作
Liudi Yang, Yang Bai, George Eskandar, Fengyi Shen, Mohammad Altillawi, Dong Chen, Ziyuan Liu, Abhinav Valada
机构
*
University of Freiburg(弗赖堡大学)
;
Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Technical University of Munich(慕尼黑技术大学)
;
Huawei Heisenberg Research Center (Munich)(华为海森堡研究中心)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
感知、对话,然后适应:面向开放世界视频识别的多模态基础模型知识迁移
Boyu Chen, Siran Chen, Kunchang Li, Qinglin Xu, Yu Qiao, Yali Wang
机构
*
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
the School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Shanghai AI Laboratory(上海人工智能实验室)
SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
SNOW:基于世界知识的时空场景理解用于开放世界具身推理
Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax
机构
*
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
Esslingen University of Applied Sciences(埃斯林根应用科学大学)
;
Dr. Ing. h.c. F. Porsche AG(德意志联邦汽车工业联合会)
;
University of Michigan(密歇根大学)
;
Voxel51 Inc.(Voxel51公司)
Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology
多模态检索增强生成在生物医学领域中的增强策略探索:一项评估糖生物学问答的案例研究
Primož Kocbek, Azra Frkatović-Hodžić, Dora Lalić, Vivian Hui, Gordan Lauc, Gregor Štiglic
机构
*
University of Maribor, Faculty of Health Sciences(莫拉维亚大学健康科学学院)
;
University of Ljubljana, Medical Factory(卢布尔雅那大学医疗工厂)
;
Genos Ltd(基因公司)
;
Center for Smart Health, School of Nursing The Hong Kong Polytechnic University(智能健康中心护理学院香港理工大学)
;
University of Zagreb, Faculty of Pharmacy(扎格雷布大学药学院)
;
Usher Institute University of Edinburgh(埃德蒙顿大学usher研究所)
机构
*
School of Xingzhi College, South China Normal University(星智学院,华南师范大学)
;
College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳科技大学)
;
School of Data Science and Engineering, Xingzhi College, South China Normal University(数据科学与工程学院,星智学院,华南师范大学)
;
Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(电子与计算机工程系,香港科技大学)
;
College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)
Radiology Report Generation with Layer-Wise Anatomical Attention
基于层间解剖注意力的放射报告生成
Emmanuel D. Muñiz-De-León, Jorge A. Rosales-de-Golferichs, Ana S. Muñoz-Rodríguez, Alejandro I. Trejo-Castro, Eduardo de Avila-Armenta, Antonio Martínez-Torteya
机构
*
School of Computer Science, China University of Geosciences (Wuhan)(中国地质大学(武汉)计算机科学学院)
;
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
School of Computer Science, Wuhan University(武汉大学计算机科学学院)
;
School of Computing, Engineering and Mathematical Sciences, La Trobe University(拉筹伯大学计算科学、工程与数学科学学院)