arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-25 至 2025-12-25 共收录 35 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 8 篇

2512.20633 2025-12-25 cs.LG cs.AI 57%

Enhancing Lung Cancer Treatment Outcome Prediction through Semantic Feature Engineering Using Large Language Models

通过大语言模型进行语义特征工程提升肺癌治疗预后预测

MunHwan Lee, Shaika Chowdhury, Xiaodi Li, Sivaraman Rajaganapathy, Eric W Klee, Ping Yang, Terence Sio, Liewei Wang, James Cerhan, Nansu NA Zong

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 本研究利用大语言模型进行语义特征工程,提升肺癌治疗预后预测的准确性与可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 4 篇

2504.11467 2025-12-25 cs.CV eess.IV 79%

A Multicore and Edge TPU-Accelerated Multimodal TinyML System for Livestock Behavior Recognition

一种多核和边缘TPU加速的多模态TinyML系统用于牲畜行为识别

Qianxue Zhang, Eiman Kanjo

机构 * Medical AI Lab, Hebei Provincial Engineering Research Center for AI-Based Cancer Treatment Decision-Making, The First Hospital of Hebei Medical University(医学人工智能实验室,河北省人工智能辅助癌症治疗决策工程研究中心,河北省医科大学第一医院) Computing Department, Imperial College London(计算部门,帝国理工学院伦敦分校) Professor Pervasive Sensing & TinyML and the Head of the Smart Sensing Lab at Nottingham Trent University(感知与TinyML教授及智能感知实验室主任,诺丁汉特伦特大学) Provost’s Visiting Professor in tinyML at Imperial College London(帝国理工学院伦敦分校副校长兼任TinyML客座教授)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种多核和边缘TPU加速的多模态TinyML系统,用于高效识别牲畜行为,实现高模型压缩和低延迟的实时推理。

Comments 12 pages, 10 figures

Journal ref IEEE Internet of Things Journal, vol. 13, no. 1, pp. 666-677, 1 Jan.1, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21200 2025-12-25 eess.SY cs.SY 78%

A Multimodal Human-Centered Framework for Assessing Pedestrian Well-Being in the Wild

一种多模态的人本框架用于评估野外行人福祉

Yasaman Hakiminejad, Arash Tavakoli

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本文提出了一种多模态人本框架,通过整合生理传感、地理跟踪和即时自我报告,评估野外行人福祉,揭示城市环境对福祉的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09141 2025-12-25 cs.RO 78%

RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation

RGMP: 基于几何先验的多模态策略用于通用人形机器人操作

Xuetao Li, Wenke Huang, Nengyuan Pan, Kaiyan Zhao, Songhua Yang, Yiming Wang, Mengde Li, Mang Ye, Jifeng Xuan, Miao Li

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 RGMP通过结合几何-语义推理与递归高斯适应,实现了高效的人形机器人多模态操作控制。

Journal ref Proceedings of the AAAI conference on artificial intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20929 2025-12-25 q-bio.NC cs.CL 57%

Decoding Predictive Inference in Visual Language Processing via Spatiotemporal Neural Coherence

通过时空神经一致性解码视觉语言处理中的预测推断

Sean C. Borneman, Julia Krebs, Ronnie B. Wilbur, Evie A. Malaia

机构 * Department of Physics(物理系) Carnegie-Mellon University(卡内基梅隆大学) Department of Linguistics(语言学系) University of Salzburg(萨尔茨堡大学) Purdue University(普渡大学) University of Alabama(阿拉巴马大学) Department of Communicative Disorders(沟通障碍系)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

AI总结 本研究通过时空神经一致性方法解码聋人手语使用者在动态视觉语言处理中的预测推断,揭示了语言理解中左侧半球和前额低频相干性的重要性,并展示了经验驱动的感知生成模型的多模态探测方法。

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Foundation Models for the Brain and Body

详情

展开后加载摘要…

URL PDF HTML 收藏