arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2807 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2807 篇

2305.15021 2023-09-15 cs.RO cs.AI cs.CV cs.LG 62%

EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Yao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang, Mingyu Ding, Jun Jin, Bin Wang, Jifeng Dai, Yu Qiao, Ping Luo

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15097 2023-08-30 cs.AI cs.CL 62%

Sequential annotations for naturally-occurring HRI: first insights

Lucien Tisserand, Frédéric Armetta, Heike Baldauf-Quilliatre, Antoine Bouquin, Salima Hassas, Mathieu Lefort

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments Peer-reviewed workshop paper accepted for the ''Human-Robot Conversational Interaction'' workshop that took place at the ''ACM/IEEE International Conference on Human-Robot Interaction'' 2023 Conference in Stockholm, Sweden

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11877 2023-08-25 cs.CV cs.AI 62%

Integrated Image and Location Analysis for Wound Classification: A Deep Learning Approach

Yash Patel, Tirth Shah, Mrinal Kanti Dhar, Taiyu Zhang, Jeffrey Niezgoda, Sandeep Gopalakrishnan, Zeyun Yu

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16207 2023-06-29 cs.AI cs.CL cs.RO 62%

Inferring the Goals of Communicating Agents from Actions and Instructions

Lance Ying, Tan Zhi-Xuan, Vikash Mansinghka, Joshua B. Tenenbaum

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

Comments 8 pages, 5 figures. Accepted to the ICML 2023 Workshop on Theory of Mind in Communicating Agents. Supplementary Information: https://osf.io/gh758/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14911 2023-06-28 cs.CL cs.AI 62%

"You might think about slightly revising the title": identifying hedges in peer-tutoring interactions

Yann Raphalen, Chloé Clavel, Justine Cassell

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments Published in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), 2022

Journal ref Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), Volume 1: long papers (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09349 2023-06-13 cs.AI cs.CL cs.RO 62%

LLM as A Robotic Brain: Unifying Egocentric Memory and Control

Jinjie Mai, Jun Chen, Bing Li, Guocheng Qian, Mohamed Elhoseiny, Bernard Ghanem

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments This early project is now integrated to: Mindstorms in Natural Language-Based Societies of Mind, arXiv:2305.17066

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06358 2023-05-12 cs.AI cs.CL 62%

Accessible Instruction-Following Agent

Kairui Zhou

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09448 2023-04-20 cs.LG cs.CL cs.CV 62%

EC^2: Emergent Communication for Embodied Control

Yao Mu, Shunyu Yao, Mingyu Ding, Ping Luo, Chuang Gan

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Published in CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09474 2023-04-05 cs.CV cs.AI cs.RO 62%

3D Object Detection for Autonomous Driving: A Comprehensive Survey

Jiageng Mao, Shaoshuai Shi, Xiaogang Wang, Hongsheng Li

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to International Journal of Computer Vision (IJCV). Project page is at https://github.com/PointsCoder/Awesome-3D-Object-Detection-for-Autonomous-Driving

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08620 2023-03-27 cs.CL cs.AI cs.CY cs.HC cs.LG 62%

POTATO: The Portable Text Annotation Tool

Jiaxin Pei, Aparna Ananthasubramaniam, Xingyao Wang, Naitian Zhou, Jackson Sargent, Apostolos Dedeloudis, David Jurgens

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2022 DEMO

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.15027 2023-01-18 cs.LG cs.AI cs.CL cs.NE 62%

Symbol Emergence as Inter-personal Categorization with Head-to-head Latent Word

Kazuma Furukawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 4 figures, 5 tables

Journal ref IEEE International Conference on Development and Learning (ICDL 2022), 2022, 60-67

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00433 2023-01-03 cs.AI cs.CV cs.IT math.IT 62%

Optimization of Image Transmission in a Cooperative Semantic Communication Networks

Wenjing Zhang, Yining Wang, Mingzhe Chen, Tao Luo, Dusit Niyato

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments 29 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08729 2022-12-20 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 62%

Distribution-aware Goal Prediction and Conformant Model-based Planning for Safe Autonomous Driving

Jonathan Francis, Bingqing Chen, Weiran Yao, Eric Nyberg, Jean Oh

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted: 1st Workshop on Safe Learning for Autonomous Driving, at the International Conference on Machine Learning (ICML 2022); Best Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.06175 2022-11-14 cs.AI cs.CL cs.LG cs.RO 62%

A Generalist Agent

Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, Nando de Freitas

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Published at TMLR, 42 pages

Journal ref Transactions on Machine Learning Research, 11/2022, https://openreview.net/forum?id=1ikK0kHjvj

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06155 2022-10-17 cs.CL cs.AI 62%

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Teng Hu, Weichong Yin, Yongfeng Chen, Yin Zhang, Shikun Feng, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2022 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02764 2022-03-08 cs.CV cs.CL cs.RO 62%

Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation

Yicong Hong, Zun Wang, Qi Wu, Stephen Gould

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.02173 2022-03-04 cs.RO cs.AI cs.CV cs.LG cs.MA 62%

Multi-Agent Variational Occlusion Inference Using People as Sensors

Masha Itkina, Ye-Ji Mun, Katherine Driggs-Campbell, Mykel J. Kochenderfer

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 9 figures, International Conference on Robotics and Automation (ICRA) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.11576 2021-11-26 cs.LG cs.CL cs.CV 62%

Building Goal-Oriented Dialogue Systems with Situated Visual Context

Sanchit Agarwal, Jan Jezabek, Arijit Biswas, Emre Barut, Shuyang Gao, Tagyoung Chung

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.00956 2021-09-02 cs.LG cs.AI cs.CL 62%

SocialAI: Benchmarking Socio-Cognitive Abilities in Deep Reinforcement Learning Agents

Grgur Kovač, Rémy Portelas, Katja Hofmann, Pierre-Yves Oudeyer

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments under review. This paper extends and generalizes work in arXiv:2104.13207

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02846 2021-08-09 cs.AI cs.CV cs.HC cs.LG cs.RO 62%

Communicative Learning with Natural Gestures for Embodied Navigation Agents with Human-in-the-Scene

Qi Wu, Cheng-Ju Wu, Yixin Zhu, Jungseock Joo

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments To appear in IROS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.13073 2021-06-24 cs.CL cs.AI 62%

Maria: A Visual Experience Powered Conversational Agent

Zujie Liang, Huang Hu, Can Xu, Chongyang Tao, Xiubo Geng, Yining Chen, Fan Liang, Daxin Jiang

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL 2021 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.10110 2021-06-21 cs.CV cs.AI cs.MA cs.RO 62%

Towards Distraction-Robust Active Visual Tracking

Fangwei Zhong, Peng Sun, Wenhan Luo, Tingyun Yan, Yizhou Wang

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.AI

Comments To appear in ICML2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.01526 2020-12-04 cs.CV cs.AI cs.RO 62%

From Goals, Waypoints & Paths To Long Term Human Trajectory Forecasting

Karttikeya Mangalam, Yang An, Harshayu Girase, Jitendra Malik

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 7 figures (including 2 GIFs)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.08744 2020-10-23 cs.CV cs.AI cs.RO 62%

PLOP: Probabilistic poLynomial Objects trajectory Planning for autonomous driving

Thibault Buhet, Emilie Wirbel, Andrei Bursuc, Xavier Perrotton

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at CorRL 2020 (matching camera-ready version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.01719 2020-10-15 cs.CL cs.AI 62%

Grounded Language Learning Fast and Slow

Felix Hill, Olivier Tieleman, Tamara von Glehn, Nathaniel Wong, Hamza Merzic, Stephen Clark

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.12394 2020-01-17 cs.CL cs.CV cs.LG 62%

All-in-One Image-Grounded Conversational Agents

Da Ju, Kurt Shuster, Y-Lan Boureau, Jason Weston

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.06315 2019-10-15 cs.CV cs.CL cs.LG 62%

Dynamic Attention Networks for Task Oriented Grounding

Soumik Dasgupta, Badri N. Patro, Vinay P. Namboodiri

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted ICCV 2019 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.02701 2019-01-11 cs.CV cs.LG cs.MM 62%

Guess What's on my Screen? Clustering Smartphone Screenshots with Active Learning

Agnese Chiatti, Dolzodmaa Davaasuren, Nilam Ram, Prasenjit Mitra, Byron Reeves, Thomas Robinson

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.MM

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.06150 2018-09-20 cs.RO cs.AI cs.CL cs.LG 62%

FollowNet: Robot Navigation by Following Natural Language Directions with Deep Reinforcement Learning

Pararth Shah, Marek Fiser, Aleksandra Faust, J. Chase Kew, Dilek Hakkani-Tur

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 8 figures

Journal ref Third Workshop in Machine Learning in the Planning and Control of Robot Motion at ICRA, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.04012 2018-06-12 cs.CV cs.MM 62%

Hierarchy of GANs for learning embodied self-awareness model

Mahdyar Ravanbakhsh, Mohamad Baydoun, Damian Campo, Pablo Marin, David Martin, Lucio Marcenaro, Carlo S. Regazzoni

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.MM

Comments 2018 IEEE International Conference on Image Processing - ICIP'18. arXiv admin note: text overlap with arXiv:1806.02609

详情

展开后加载摘要…

URL PDF HTML 收藏