arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2821 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2821 篇

2412.18354 2024-12-25 cs.AI q-bio.NC 57%

The Thousand Brains Project: A New Paradigm for Sensorimotor Intelligence

Viviane Clay, Niels Leadholm, Jeff Hawkins

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10402 2024-12-17 cs.AI cs.RO 57%

TANGO: Training-free Embodied AI Agents for Open-world Tasks

Filippo Ziliotto, Tommaso Campari, Luciano Serafini, Lamberto Ballan

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06740 2024-12-06 cs.LG cs.AI 57%

Dockformer: A transformer-based molecular docking paradigm for large-scale virtual screening

Zhangfan Yang, Junkai Ji, Shan He, Jianqiang Li, Tiantian He, Ruibin Bai, Zexuan Zhu, Yew Soon Ong

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16578 2024-12-04 cs.RO cs.AI 57%

QuadrupedGPT: Towards a Versatile Quadruped Agent in Open-ended Worlds

Yuting Mei, Ye Wang, Sipeng Zheng, Qin Jin

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17255 2024-12-03 cs.LG cs.AI 57%

APT: Architectural Planning and Text-to-Blueprint Construction Using Large Language Models for Open-World Agents

Jun Yu Chen, Tao Gao

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01537 2024-12-02 cs.CV cs.RO 57%

SceneMotion: From Agent-Centric Embeddings to Scene-Wide Forecasts

Royden Wagner, Ömer Sahin Tas, Marlon Steiner, Fabian Konstantinidis, Hendrik Königshof, Marvin Klemp, Carlos Fernandez, Christoph Stiller

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments ITSC'24; updated table VI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18073 2024-11-28 cs.AI cs.IR 57%

DuMapper: Towards Automatic Verification of Large-Scale POIs with Street Views at Baidu Maps

Miao Fan, Jizhou Huang, Haifeng Wang

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12676 2024-11-20 cs.CV cs.LG 57%

IoT-Based 3D Pose Estimation and Motion Optimization for Athletes: Application of C3D and OpenPose

Fei Ren, Chao Ren, Tianyi Lyu

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00248 2024-11-20 cs.CL 57%

A Demonstration of Adaptive Collaboration of Large Language Models for Medical Decision-Making

Yubin Kim, Chanwoo Park, Hyewon Jeong, Cristina Grau-Vilchez, Yik Siu Chan, Xuhai Xu, Daniel McDuff, Hyeonhoon Lee, Cynthia Breazeal, Hae Won Park

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL

Comments Under Review for ML4H 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10983 2024-11-19 cs.CV 57%

Framework for developing and evaluating ethical collaboration between expert and machine

Ayan Banerjee, Payal Kamboj, Sandeep Gupta

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments Accepted in ECAI Workshop AIEB

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15792 2024-11-19 cs.IR cs.AI 57%

IQLS: Framework for leveraging Metadata to enable Large Language Model based queries to complex, versatile Data

Sami Azirar, Hossam A. Gabbar, Chaouki Regoui

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08899 2024-11-15 q-fin.TR cs.AI 57%

FinVision: A Multi-Agent Framework for Stock Market Prediction

Sorouralsadat Fatemi, Yuheng Hu

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments Accepted at ICAIF 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05383 2024-11-11 cs.CL 57%

Towards Low-Resource Harmful Meme Detection with LMM Agents

Jianzhao Huang, Hongzhan Lin, Ziyan Liu, Ziyang Luo, Guang Chen, Jing Ma

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07706 2024-10-28 cs.RO cs.AI 57%

Pixel State Value Network for Combined Prediction and Planning in Interactive Environments

Sascha Rosbach, Stefan M. Leupold, Simon Großjohann, Stefan Roth

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06750 2024-10-21 cs.CY cs.AI 57%

Frontier AI Ethics: Anticipating and Evaluating the Societal Impacts of Language Model Agents

Seth Lazar

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13768 2024-10-18 cond-mat.mtrl-sci cond-mat.dis-nn cond-mat.mes-hall cs.AI cs.MA 57%

Rapid and Automated Alloy Design with Graph Neural Network-Powered LLM-Driven Multi-Agent Systems

Alireza Ghafarollahi, Markus J. Buehler

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05982 2024-10-10 cs.CV cs.RO 57%

DeMo: Decoupling Motion Forecasting into Directional Intentions and Dynamic States

Bozhou Zhang, Nan Song, Li Zhang

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03907 2024-10-08 cs.CL 57%

ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities

Ying Su, Zhan Ling, Haochen Shi, Jiayang Cheng, Yauwai Yim, Yangqiu Song

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL

Comments 13 pages, 9 figures, 8 tables, accepted to EMNLP 2024 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03303 2024-10-07 cs.LG cs.CV 57%

SELU: Self-Learning Embodied MLLMs in Unknown Environments

Boyu Li, Haobin Jiang, Ziluo Ding, Xinrun Xu, Haoran Li, Dongbin Zhao, Zongqing Lu

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18475 2024-09-30 cs.AI cs.HC 57%

Data Analysis in the Era of Generative AI

Jeevana Priya Inala, Chenglong Wang, Steven Drucker, Gonzalo Ramos, Victor Dibia, Nathalie Riche, Dave Brown, Dan Marshall, Jianfeng Gao

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02374 2024-09-26 cs.CL 57%

Conversational Health Agents: A Personalized LLM-Powered Agent Framework

Mahyar Abbasian, Iman Azimi, Amir M. Rahmani, Ramesh Jain

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments 23 pages, 6 figures, 2 tables, 4 appendices, journal paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02522 2024-09-24 cs.AI cs.RO 57%

Cog-GA: A Large Language Models-based Generative Agent for Vision-Language Navigation in Continuous Environments

Zhiyuan Li, Yanfeng Lu, Yao Mu, Hong Qiao

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13050 2024-09-24 cs.HC cs.AI 57%

Human-Centered LLM-Agent User Interface: A Position Paper

Daniel Chin, Yuxuan Wang, Gus Xia

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12889 2024-09-24 cs.AI 57%

Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case

Peng Chen, Pi Bu, Jun Song, Yuan Gao, Bo Zheng

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12005 2024-09-20 cs.RO cs.AI 57%

Representing Positional Information in Generative World Models for Object Manipulation

Stefano Ferraro, Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Sai Rajeswar

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12167 2024-09-19 eess.IV cs.CV 57%

multiPI-TransBTS: A Multi-Path Learning Framework for Brain Tumor Image Segmentation Based on Multi-Physical Information

Hongjun Zhu, Jiaohang Huang, Kuo Chen, Xuehui Ying, Ying Qian

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07129 2024-09-12 cs.CV 57%

MVLLaVA: An Intelligent Agent for Unified and Flexible Novel View Synthesis

Hanyu Jiang, Jian Xue, Xing Lan, Guohong Hu, Ke Lu

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments project page: https://jamesjg.github.io/MVLLaVA_homepage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03515 2024-09-10 cs.RO cs.AI 57%

A Study on Prompt Injection Attack Against LLM-Integrated Mobile Robotic Systems

Wenxiao Zhang, Xiangrui Kong, Conan Dewitt, Thomas Braunl, Jin B. Hong

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01630 2024-09-04 cs.RO cs.AI cs.ET 57%

SafeEmbodAI: a Safety Framework for Mobile Robots in Embodied AI Systems

Wenxiao Zhang, Xiangrui Kong, Thomas Braunl, Jin B. Hong

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20091 2024-09-04 cs.CV cs.HC cs.LG 57%

VAAD: Visual Attention Analysis Dashboard applied to e-Learning

Miriam Navarro, Álvaro Becerra, Roberto Daza, Ruth Cobos, Aythami Morales, Julian Fierrez

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Published in IEEE Intl. Symposium on Computers in Education (SIIE) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏