arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-05 至 2025-09-05 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2509.04324 2025-09-05 cs.RO cs.CV 79%

OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection

Chen Hu, Shan Luo, Letizia Gionfrida

机构 * Department of Informatics, King's College London(伦敦国王学院信息学院) Department of Engineering, King's College London(伦敦国王学院工程学院) John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛大学约翰·A·保罗森工程与应用科学学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14904 2025-09-05 cs.CV cs.AI 73%

TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP

Fan Li, Zanyi Wang, Zeyi Huang, Guang Dai, Jingdong Wang, Mengmeng Wang

机构 * Xi’an Jiaotong University(西安交通大学) SGIT AI Lab(SGIT人工智能实验室) Zhejiang University of Technology(浙江工业大学) Huawei(华为)

专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03800 2025-09-05 cs.CV 57%

MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting

Yuheng Li, Yenho Chen, Yuxiang Lai, Jike Zhong, Vanessa Wildman, Xiaofeng Yang

机构 * Department of Biomedical Engineering(生物医学工程系) Georgia Institute of Technology(佐治亚理工学院) Department of Machine Learning(机器学习系) Department of Radiation Oncology(放射肿瘤科) Emory University School of Medicine(埃默里大学医学院) University of Southern California(南加州大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04162 2025-09-05 cs.AR 50%

Real Time FPGA Based Transformers & VLMs for Vision Tasks: SOTA Designs and Optimizations

Safa Mohammed Sali, Mahmoud Meribout, Ashiyana Abdul Majeed

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏