arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2821 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2821 篇

2309.10375 2023-09-20 cs.CV 57%

Pointing out Human Answer Mistakes in a Goal-Oriented Visual Dialogue

Ryosuke Oshima, Seitaro Shinagawa, Hideki Tsunashima, Qi Feng, Shigeo Morishima

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Accepted at ICCVW 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00923 2023-09-19 cs.RO cs.CV cs.HC 57%

Sonicverse: A Multisensory Simulation Platform for Embodied Household Agents that See and Hear

Ruohan Gao, Hao Li, Gokul Dharan, Zhuzhu Wang, Chengshu Li, Fei Xia, Silvio Savarese, Li Fei-Fei, Jiajun Wu

专题命中 多模态Agent :audio-visual(abstract);分类 cs.CV

Comments In ICRA 2023. Project page: https://ai.stanford.edu/~rhgao/sonicverse/. Code: https://github.com/StanfordVL/sonicverse. Gao and Li contributed equally to this work and are in alphabetical order

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05036 2023-09-12 cs.RO cs.CV 57%

What Is Near?: Room Locality Learning for Enhanced Robot Vision-Language-Navigation in Indoor Living Environments

Muraleekrishna Gopinathan, Jumana Abu-Khalaf, David Suter, Sidike Paheding, Nathir A. Rawashdeh

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01073 2023-09-06 cs.CV 57%

Spatial and Visual Perspective-Taking via View Rotation and Relation Reasoning for Embodied Reference Understanding

Cheng Shi, Sibei Yang

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments ECCV 2022. Code: http://github.com/ChengShiest/REP-ERU

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10324 2023-09-01 cs.RO cs.AI cs.HC 57%

HARPS: An Online POMDP Framework for Human-Assisted Robotic Planning and Sensing

Luke Burks, Hunter M. Ray, Jamison McGinley, Sousheel Vunnam, Nisar Ahmed

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to IEEE Transactions on Robotics. 20 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09066 2023-08-21 cs.CV 57%

PatchCT: Aligning Patch Set and Label Set with Conditional Transport for Multi-Label Image Classification

Miaoge Li, Dongsheng Wang, Xinyang Liu, Zequn Zeng, Ruiying Lu, Bo Chen, Mingyuan Zhou

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

Comments accepted by ICCV23

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07751 2023-08-16 cs.CV 57%

CASPNet++: Joint Multi-Agent Motion Prediction

Maximilian Schäfer, Kun Zhao, Anton Kummert

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06498 2023-08-15 cs.AI cs.HC cs.RO 57%

Latent Emission-Augmented Perspective-Taking (LEAPT) for Human-Robot Interaction

Kaiqi Chen, Jing Yu Lim, Kingsley Kuan, Harold Soh

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.05221 2023-08-11 cs.HC cs.AI cs.RO 57%

Alexa, play with robot: Introducing the First Alexa Prize SimBot Challenge on Embodied AI

Hangjie Shi, Leslie Ball, Govind Thattai, Desheng Zhang, Lucy Hu, Qiaozi Gao, Suhaila Shakiah, Xiaofeng Gao, Aishwarya Padmakumar, Bofei Yang, Cadence Chung, Dinakar Guthy, Gaurav Sukhatme, Karthika Arumugam, Matthew Wen, Osman Ipek, Patrick Lange, Rohan Khanna, Shreyas Pansare, Vasu Sharma, Chao Zhang, Cris Flagg, Daniel Pressel, Lavina Vaz, Luke Dai, Prasoon Goyal, Sattvik Sahai, Shaohua Liu, Yao Lu, Anna Gottardi, Shui Hu, Yang Liu, Dilek Hakkani-Tur, Kate Bland, Heather Rocker, James Jeun, Yadunandana Rao, Michael Johnston, Akshaya Iyengar, Arindam Mandal, Prem Natarajan, Reza Ghanadan

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.02139 2023-08-03 cs.AI cs.RO 57%

Data Association Aware POMDP Planning with Hypothesis Pruning Performance Guarantees

Moran Barenboim, Idan Lev-Yehudi, Vadim Indelman

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00291 2023-08-02 eess.IV cs.CV 57%

Fundus-Enhanced Disease-Aware Distillation Model for Retinal Disease Classification from OCT Images

Lehan Wang, Weihang Dai, Mei Jin, Chubin Ou, Xiaomeng Li

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments Accepted as a conference paper at MICCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10810 2023-07-21 cs.LG cs.AI 57%

On Combining Expert Demonstrations in Imitation Learning via Optimal Transport

Ilana Sebag, Samuel Cohen, Marc Peter Deisenroth

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Journal ref NeurIPS Workshop on Optimal Transport and Machine Learning, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02280 2023-07-06 cs.CV 57%

Interactive Image Segmentation with Cross-Modality Vision Transformers

Kun Li, George Vosselman, Michael Ying Yang

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12071 2023-06-30 cs.CV cs.RO 57%

ProphNet: Efficient Agent-Centric Motion Forecasting with Anchor-Informed Proposals

Xishun Wang, Tong Su, Fang Da, Xiaodong Yang

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments CVPR 2023 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12387 2023-06-22 cs.CL cs.HC 57%

Solving Dialogue Grounding Embodied Task in a Simulated Environment using Further Masked Language Modeling

Weijie Jack Zhang

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.07857 2023-06-21 cs.RO cs.CV 57%

GATraj: A Graph- and Attention-based Multi-Agent Trajectory Prediction Model

Hao Cheng, Mengmeng Liu, Lin Chen, Hellward Broszio, Monika Sester, Michael Ying Yang

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18898 2023-05-31 cs.RO cs.AI 57%

AlphaBlock: Embodied Finetuning for Vision-Language Reasoning in Robot Manipulation

Chuhao Jin, Wenhui Tan, Jiange Yang, Bei Liu, Ruihua Song, Limin Wang, Jianlong Fu

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13076 2023-05-23 cs.CL 57%

An Abstract Specification of VoxML as an Annotation Language

Kiyong Lee, Nikhil Krishnaswamy, James Pustejovsky

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments 8 pages, 4 figures, Proceedings of 19th Joint ISO-ACL Workshop on Interoperable Semantic Annotation (ISA 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10783 2023-05-19 cs.AI 57%

Transforming Human-Centered AI Collaboration: Redefining Embodied Agents Capabilities through Interactive Grounded Language Instructions

Shrestha Mohanty, Negar Arabzadeh, Julia Kiseleva, Artem Zholus, Milagro Teruel, Ahmed Awadallah, Yuxuan Sun, Kavya Srinet, Arthur Szlam

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09600 2023-05-17 cs.AI cs.LG 57%

Deep Reinforcement Learning to Maximize Arterial Usage during Extreme Congestion

Ashutosh Dutta, Milan Jain, Arif Khan, Arun Sathanur

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07735 2023-05-04 cs.AI cs.RO 57%

Monte Carlo Planning in Hybrid Belief POMDPs

Moran Barenboim, Moshe Shienman, Vadim Indelman

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08562 2023-03-16 cs.CV 57%

MGA: Medical generalist agent through text-guided knowledge transformation

Weijian Huang, Hao Yang, Cheng Li, Mingtong Dai, Rui Yang, Shanshan Wang

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.04129 2023-03-08 cs.AI cs.LG 57%

Foundation Models for Decision Making: Problems, Methods, and Opportunities

Sherry Yang, Ofir Nachum, Yilun Du, Jason Wei, Pieter Abbeel, Dale Schuurmans

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08590 2023-02-20 cs.CL 57%

What A Situated Language-Using Agent Must be Able to Do: A Top-Down Analysis

David Schlangen

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03062 2022-11-30 eess.IV cs.CV 57%

MyoPS-Net: Myocardial Pathology Segmentation with Flexible Combination of Multi-Sequence CMR Images

Junyi Qiu, Lei Li, Sihan Wang, Ke Zhang, Yinyin Chen, Shan Yang, Xiahai Zhuang

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03049 2022-11-08 cs.RO cs.AI q-bio.NC 57%

Learning body models: from humans to humanoids

Matej Hoffmann

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 34 pages, 5 figures. Habilitation thesis, Faculty of Electrical Engineering, Czech Technical University in Prague (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.09325 2022-11-08 q-bio.NC cs.AI 57%

Body models in humans, animals, and robots: mechanisms and plasticity

Matej Hoffmann

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 27 pages, 8 figures

Journal ref 2021, Body Schema and Body Image: New Directions, Oxford University Press, pp. 152-180

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.11746 2022-09-27 cs.AI 57%

Evaluating Agent Interactions Through Episodic Knowledge Graphs

Selene Báez Santamaría, Piek Vossen, Thomas Baier

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to 1st Workshop on Customized Chat Grounding Persona and Knowledge, at COLING (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.10126 2022-09-22 cs.CV cs.LG 57%

Exploring Modulated Detection Transformer as a Tool for Action Recognition in Videos

Tomás Crisol, Joel Ermantraut, Adrián Rostagno, Santiago L. Aggio, Javier Iparraguirre

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments 5 pages, 2 figures, 1 results chart

Journal ref JAIIO - JORNADAS ARGENTINAS DE INFORMATICA 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.06739 2022-08-16 eess.IV cs.CV cs.LG physics.med-ph 57%

Machine Learning Based Radiomics for Glial Tumor Classification and Comparison with Volumetric Analysis

Sevcan Turk, Kaya Oguz, Mehmet Orman, Emre Caliskan, Yesim Ertan, Erkin Ozgiray, Taner Akalin, Ashok Srinivasan, Omer Kitis

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏