arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-05 至 2025-09-05 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 5 篇

2509.04254 2025-09-05 cs.HC 82%

MuMTAffect: A Multimodal Multitask Affective Framework for Personality and Emotion Recognition from Physiological Signals

Meisam Jamshidi Seikavandi, Fabricio Batista Narcizo, Ted Vucurevich, Andrew Burke Dittberner, Paolo Burelli

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04330 2025-09-05 cs.IR 78%

Temporal Interest-Driven Multimodal Personalized Content Generation

Tian Miao

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04210 2025-09-05 cs.CE cs.LG 78%

COBRA: Multimodal Sensing Deep Learning Framework for Remote Chronic Obesity Management via Wrist-Worn Activity Monitoring

Zhengyang Shen, Bo Gao, Mayue Shi

机构 * Department of Electrical and Electronic Engineering, Imperial College London, London SW7 2AZ, UK(帝国理工学院电子与电气工程系) Institute of Biomedical Engineering, Department of Engineering Science, University of Oxford, Oxford OX3 7DQ, UK(牛津大学生物医学工程研究所)

专题命中 视频多模态 :multimodal(title,abstract)

Comments 19 pages, 4 figures. *Correspondence: m.shi16@imperial.ac.uk. Accepted by the IUPESM World Congress on Medical Physics and Biomedical Engineering 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19346 2025-09-05 cs.LG 78%

Short-Form Video Recommendations with Multimodal Embeddings: Addressing Cold-Start and Bias Challenges

Andrii Dzhoha, Katya Mirylenka, Egor Malykh, Marco-Andrea Buchmann, Francesca Catino

机构 * Zalando SE Berlin Germany(泽尔安多德国分公司) Zalando Switzerland AG Zürich Switzerland(泽尔安多瑞士分公司)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04117 2025-09-05 cs.CV 57%

DVS-PedX: Synthetic-and-Real Event-Based Pedestrian Dataset

Mustafa Sakhai, Kaung Sithu, Min Khant Soe Oke, Maciej Wielgosz

机构 * Faculty of Computer Science, Electronics and Telecommunications(计算机科学与电子技术学院) AGH University of Science and Technology(AGH科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 12 pages, 8 figures, 3 tables; dataset descriptor paper introducing DVS-PedX (synthetic-and-real event-based pedestrian dataset with baselines) External URL: https://doi.org/10.5281/zenodo.17030898

详情

展开后加载摘要…

URL PDF HTML 收藏