arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9204 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9204 篇

2606.22864 2026-06-23 cs.LG 新提交 79%

When AUC 0.998 Is Not Enough: A Candidate Evaluation Protocol for Hidden-State Probes of Indirect Prompt Injection in Multimodal Computer-Use Agents

当AUC 0.998还不够:多模态计算机使用智能体中隐藏状态探针对间接提示注入的候选评估协议

Yanhang Li, Zhichao Fan, Zexin Zhuang

机构 * Northeastern University(东北大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Southern Methodist University(南卫理公会大学)

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本文通过单骨干案例研究,论证高AUC不能直接证明恶意内容检测,提出后验诊断和候选控制集来明确探针能力的边界。

Comments 17 pages, 3 figures. Camera-ready version for EvalMG '26, The 2nd Workshop on Evaluation for Multimodal Generation, co-located with SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19604 2025-09-25 cs.LG 79%

Improved Therapeutic Antibody Reformatting through Multimodal Machine Learning

Jiayi Xin, Aniruddh Raghu, Nick Bhattacharya, Adam Carr, Melanie Montgomery, Hunter Elliott

机构 * University of Pennsylvania(宾夕法尼亚大学) BigHat Biosciences(BigHat生物技术公司)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(comments)

Comments NeurIPS 2025 AI4Science Workshop and NeurIPS 2025 Multi-modal Foundation Models and Large Language Models for Life Sciences Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03543 2025-05-07 cs.IR 79%

1$^{st}$ Place Solution of WWW 2025 EReL@MIR Workshop Multimodal CTR Prediction Challenge

Junwei Xu, Zehao Zhao, Xiaoyu Hu, Zhenjie Song

专题命中 多模态评测 :multimodal(title,abstract)

Comments Technical report for the 1$^{st}$ place solution of WWW 2025 EReL@MIR Workshop Multimodal CTR Prediction Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04247 2024-07-08 cs.CL cs.AI cs.CV 79%

ArAIEval Shared Task: Propagandistic Techniques Detection in Unimodal and Multimodal Arabic Content

Maram Hasanain, Md. Arid Hasan, Fatema Ahmed, Reem Suwaileh, Md. Rafiul Biswas, Wajdi Zaghouani, Firoj Alam

专题命中 多模态评测 :multimodal(title,comments);分类 cs.CV、cs.CL、cs.AI

Comments propaganda, span detection, disinformation, misinformation, fake news, LLMs, GPT-4, multimodality, multimodal LLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00709 2023-10-06 cs.HC cs.CY 79%

Understanding the Social Context of Eating with Multimodal Smartphone Sensing: The Role of Country Diversity

Nathan Kammoun, Lakmal Meegahapola, Daniel Gatica-Perez

专题命中 多模态评测 :multimodal(title,abstract)

Comments 25th ACM International Conference on Multimodal Interaction (ICMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.07526 2023-01-19 cs.LG 79%

AutoFraudNet: A Multimodal Network to Detect Fraud in the Auto Insurance Industry

Azin Asgarian, Rohit Saha, Daniel Jakubovitz, Julia Peyre

专题命中 多模态评测 :multimodal(title,abstract)

Comments Published at The AAAI-2023 Workshop On Multimodal AI For Financial Forecasting

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.07519 2022-10-14 eess.SP cs.IT math.IT 79%

Multi-Modal Beam Prediction Challenge 2022: Towards Generalization

Gouranga Charan, Umut Demirhan, João Morais, Arash Behboodi, Hamed Pezeshki, Ahmed Alkhateeb

专题命中 多模态评测 :multi-modal(title,abstract)

Comments The dataset is available on the ML competition page: https://deepsense6g.net/multi-modal-beam-prediction-challenge/

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.10214 2018-02-01 cs.RO 79%

CREATE: Multimodal Dataset for Unsupervised Learning, Generative Modeling and Prediction of Sensory Data from a Mobile Robot in Indoor Environments

Simon Brodeur, Simon Carrier, Jean Rouat

专题命中 多模态评测 :multimodal(title,abstract)

Comments The CREATE dataset is Open access and available on IEEE Dataport (https://ieee-dataport.org/open-access/create-multimodal-dataset-unsupervised-learning-and-generative-modeling-sensory-data)

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.03152 2015-11-13 cs.RO 79%

A Handheld Device for the In Situ Acquisition of Multimodal Tactile Sensing Data

Joshua Wade, Tapomayukh Bhattacharjee, Charles C. Kemp

专题命中 多模态评测 :multimodal(title,abstract)

Comments This short 2-page paper was accepted and presented as a poster (https://goo.gl/8TqDTK) on Sep. 28, 2015 in IROS 2015 Workshop on 'See and Touch: 1st Workshop on multimodal sensor-based robot control for HRI and soft manipulation" organized by A. Cherubini, Y. Mezouar, D. Navarro-Alarcon, M. Prats, and J. A. Corrales Ramon. It was peer reviewed by 2 reviewers

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07543 2026-08-11 cs.CV cs.AI 新提交 79%

Performance of large language models in the optical diagnosis of colorectal polyps

大型语言模型在结直肠息肉光学诊断中的性能

Joshua C. Vences, William T. Tran, Nikko Gimpaya, Catharine M. Walsh, Rishad J. Khan, Robert Bechara, Asher C. Wiggins, Celine N. Rousan, Kaitlyn V. G. L. Morgado, Angie Ibrahim, Kevin H. M. Kuo, Daniel von Renteln, Alexander Hann, Dennis L. Shung, Michael A. Scaffidi, Charles Ménard, Joshua Landy, Samir C. Grover

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本研究评估Claude Opus 4等5款大型语言模型对结直肠息肉的光学诊断性能,发现其区分息肉亚型的准确率接近专家共识,但灵敏度与特异度未达ESGE标准,需进一步研究方可临床应用。

Comments 22 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06420 2026-08-07 cs.CV cs.AI 交叉投稿 79%

HoloCount: A Holistic Visual Counting Benchmark for MLLMs

HoloCount:用于多模态大语言模型的整体视觉计数基准

Jinhong Deng, Limeng Qiao, Guanglu Wan

机构 * Meituan(美团)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 研究针对多模态大语言模型定量精度瓶颈及现有计数基准局限,引入HoloCount基准,从语义、分析计数和鲁棒性测试三方面评估模型,经对20多个模型详尽评估,揭示其性能差距,为多模态系统发展提供路线图。

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06276 2026-07-29 cs.CV cs.AI 版本更新 79%

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension

RefBench-PRO:面向感知与推理的指称表达理解基准

Tianyi Gao, Hao Li, Han Fang, Xin Wei, Xiaodong Dong, Hongbo Sun, Ye Yuan, Zhongjiang He, Jinglin Xu, Jingmin Xin, Hao Sun

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(国家人机混合增强智能重点实验室) National Engineering Research Center for Visual Information and Applications(国家视觉信息与应用工程技术研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) Xi’an Jiaotong University(西安交通大学) Institute of Artificial Intelligence (TeleAI)(人工智能研究院(TeleAI)) China Telecom(中国电信) Shanghai Jiao Tong University(上海交通大学) University of Science and Technology Beijing(北京科技大学)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 RefBench-PRO提出一个面向感知与推理的指称表达理解基准,通过分解指称表达为感知和推理两个维度,设计六个逐步递增的挑战任务,并引入Ref-R1强化学习方案提升定位精度,实现对多模态大语言模型的可解释评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14145 2026-06-23 cs.CL cs.CV 版本更新 79%

MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos

MMOU:面向长且复杂真实世界视频的大规模多任务全模态理解与推理基准

Arushi Goel, Sreyan Ghosh, Vatsal Agarwal, Nishit Anand, Kaousheik Jayakumar, Lasha Koroshinadze, Yao Xu, Katie Lyons, James Case, Karan Sapra, Kevin J. Shih, Siddharth Gururani, Abhinav Shrivastava, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Bryan Catanzaro, Mohammad Shoeybi, Wei Ping

机构 * NVIDIA, USA(NVIDIA美国分公司) University of Maryland, College Park, USA(马里兰大学帕克分校)

专题命中 多模态评测 :multimodal(abstract);audio-visual(abstract);omni-modal(abstract);分类 cs.CV、cs.CL

AI总结 提出MMOU基准,包含2万个人工标注问题与11877个长视频,覆盖13项技能,评估多模态大模型在长视频中的全模态理解与推理能力,最佳闭源模型仅64.2%准确率,开源模型46.8%。

Comments Project Page: https://huggingface.co/datasets/nvidia/MMOU

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18621 2026-05-19 cs.CV cs.AI 79%

CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark

CrossView Suite: 利用数据集、模型和基准 harnessing MLLMs 的跨视图空间智能

Wei Wang, Yuqian Yuan, Tianwei Lin, Wenqiao Zhang, Siliang Tang, Jun Xiao, Yueting Zhuang

机构 * Zhejiang University(浙江大学)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 该研究提出CrossView Suite,通过开发CrossViewSet、CrossViewBench和CrossViewer三个组件,解决跨视图推理中的数据稀缺、评估不足和对齐机制缺失问题,提升多视图空间理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13169 2026-05-18 cs.CV cs.AI 79%

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World

PanoWorld:迈向360度全景世界的空间超感知

Changpeng Wang, Xin Lin, Junhan Liu, Yuheng Liu, Zhen Wang, Donglian Qi, Yunfeng Yan, Xi Chen

机构 * Zhejiang University(浙江大学) University of California, San Diego(加州大学圣地亚哥分校) University of California, Irvine(加州大学伊维特分校) The University of Hong Kong(香港大学)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出PanoWorld,通过构建全景原生理解能力,解决传统多模态大模型在空间感知上的不足,通过全景空间交叉注意力机制提升3D空间推理能力,并建立PanoSpace-Bench基准测试,验证了全景原生监督的有效性。

Comments Project page: https://wcpcp.github.io/PanoWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13530 2026-05-14 cs.CV cs.AI 79%

Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs

迈向统一的手术场景理解:通过多模态大语言模型弥合推理与 grounding

Jincai Huang, Shihao Zou, Yuchen Guo, Jingjing Li, Wei Ji, Kai Wang, Shanshan Wang, Weixin Si

机构 * Southern University of Science and Technology(南方科技大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Northwestern University(西北大学) University of Alberta(阿尔伯塔大学) Yale University(耶鲁大学) Nanfang Hospital(南华医院) Shenzhen University of Advanced Technology(深圳大学先进技术研究院)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出SurgMLLM框架,通过统一推理与视觉 grounding 实现手术场景理解,提升三元组识别和分割精度,实验表明其在三元组识别指标AP_IVT上提升显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12882 2026-05-14 cs.CL cs.CV 79%

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

CiteVQA:可信赖文档智能的证据归因基准测试

Dongsheng Ma, Jiayu Li, Zhengren Wang, Yijie Wang, Jiahao Kong, Weijun Zeng, Jutao Xiao, Jie Yang, Wentao Zhang, Bin Wang, Conghui He

机构 * Peking University(北京大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL

AI总结 CiteVQA通过要求模型返回元素级边界框引用,评估文档理解中的证据归因,揭示了现有评估方法的可靠性缺口,主要贡献是引入Strict Attributed Accuracy指标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07141 2026-05-11 cs.CV cs.AI 79%

Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding

Qwen3-VL-Seg: 解锁基于视觉-语言接地的开放世界指代分割

Yuan Yao, Qiushi Yang, Humen Zhong, Jiangning Wei, Yifang Men, Shuai Bai, Miaomiao Cui, Zhibo Yang

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Qwen3-VL-Seg通过视觉-语言接地框架实现开放世界指代分割,采用轻量级的框引导掩码解码器,结合多尺度空间特征注入和迭代掩码感知查询优化,实现参数高效且精确的像素级分割。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00979 2026-04-23 cs.CV cs.AI 79%

IVY-FAKE: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection

IVY-FAKE:图像和视频AIGC检测的统一可解释框架和基准

Changjiang Jiang, Wenhui Dong, Zhonghao Zhang, Fengchang Yu, Wei Peng, Xinbin Yuan, Yifei Bi, Ming Zhao, Zian Zhou, Chenyang Si, Caifeng Shan

机构 * Nanjing University+(南京大学+)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出IVY-FAKE,首个大规模多模态可解释AIGC检测基准,包含106000+丰富标注样本和5000个手动验证示例,通过GRPO强化学习模型实现可解释推理,提升多基准检测性能,显著优于现有方法。

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19685 2026-04-21 cs.CV cs.AI 79%

Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline

生成篡改面部图像的归因报告:一个数据集和基线

Jingchun Lian, Lingyu Liu, Yaxiong Wang, Yujiao Wu, Lianwei Wu, Li Zhu, Zhedong Zheng

机构 * Xi’an Jiaotong University(西安交通大学) Hefei University of Technology(合肥工业大学) CSIRO(澳大利亚联邦科学与工业研究组织) Northwestern Polytechnical University(西北工业大学) University of Macau(澳门大学)

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出多模态篡改追踪数据集和 ForgeryTalker 框架,通过联合定位篡改区域和生成自然语言解释,提升面部伪造检测的语义理解能力。

Comments Accepted to ACL 2026 (Main Conference). This version includes camera-ready revisions and updated experimental results

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15871 2026-04-20 cs.CV cs.AI 79%

UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs

UniEditBench: 一种统一且成本效益高的图像和视频编辑基准,通过压缩的多模态大语言模型

Lifan Jiang, Tianrun Wu, Yuhang Pei, Chenyang Wang, Boxi Wu, Deng Cai

机构 * Zhejiang University(浙江大学)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 UniEditBench提出统一的图像和视频编辑基准,支持基于重建和指令驱动的方法,通过压缩的多模态大语言模型提供多维评分,减少部署成本并提高评估效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13756 2026-04-16 cs.CL cs.CV 79%

MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging

MedRCube:面向医学影像中多模态大语言模型细粒度与深入评估的多维框架

Zhijie Bao, Fangke Chen, Licheng Bao, Chenhui Zhang, Wei Chen, Jiajie Peng, Zhongyu Wei

机构 * School of Data Science, Fudan University(复旦大学数据科学学院) Shanghai Innovation Institute(上海创新研究院) School of Integrated Circuits, Zhejiang University(浙江大学集成电路学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院)

专题命中 多模态评测 :MLLM(summary_cn);multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出MedRCube框架,通过两阶段构建流程对33个MLLM进行细粒度评估,揭示先前方法无法获取的洞察,并引入可信度评估子集,发现快捷行为与诊断性能的显著正相关,引发临床部署的担忧。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06908 2025-11-25 cs.CV cs.AI 79%

Find Them All: Unveiling MLLMs for Versatile Person Re-identification

找到它们全部:揭示MLLMs用于多功能人物重识别

Jinhao Li, Zijian Chen, Lirong Deng, Guangtao Zhai, Changbo Wang

机构 * School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院) Institute of Image Communication and Information Processing, Shanghai Jiao Tong University(上海交通大学图像通信与信息处理研究所) Macao Polytechnic University(澳门 polytechnic 大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态评测 :multi-modal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出VP-ReID基准,利用MLLMs提升人物重识别的多功能性和有效性,同时揭示其在处理某些模态时的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15867 2025-09-04 cs.CV cs.AI 79%

TruthLens: Visual Grounding for Universal DeepFake Reasoning

Rohit Kundu, Shan Jia, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07687 2025-05-14 cs.CV cs.MM 79%

FMNV: A Dataset of Media-Published News Videos for Fake News Detection

Yihao Wang, Zhong Qian, Peifeng Li

机构 * Yihao Wang(王毅浩) Zhong Qian(钱中) Peifeng Li(李培峰)

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08002 2025-04-03 cs.CL cs.CV 79%

Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation

Juan Manuel Zambrano Chaves, Shih-Cheng Huang, Yanbo Xu, Hanwen Xu, Naoto Usuyama, Sheng Zhang, Fei Wang, Yujia Xie, Mahmoud Khademi, Ziyi Yang, Hany Awadalla, Julia Gong, Houdong Hu, Jianwei Yang, Chunyuan Li, Jianfeng Gao, Yu Gu, Cliff Wong, Mu Wei, Tristan Naumann, Muhao Chen, Matthew P. Lungren, Akshay Chaudhari, Serena Yeung-Levy, Curtis P. Langlotz, Sheng Wang, Hoifung Poon

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Journal ref Nature Communications volume 16, Article number: 3108 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25316 2026-08-27 cs.HC 新提交 78%

AVI-Personality: A Trait-Activated Multimodal Dataset for Personality and Competency Assessment in Asynchronous Video Interviews

AVI-Personality:用于异步视频面试中人格与能力评估的特质激活多模态数据集

Tianyi Zhang, Jinwenxi Shang, Antonis Koutsoumpis, Yuan Zong, Reinout E. de Vries, Wenming Zheng

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本文提出AVI-Personality数据集,包含646名参与者的3876段异步视频面试,经多维度验证,可用于开发评估人格与能力的AI模型,多模态方法性能最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25126 2026-08-27 cs.LG 新提交 78%

Multimodal Injury Risk Prediction in Tennis

网球中的多模态损伤风险预测

Francisco Erramuspe Alvarez, Shobharani Polasa, Weihao Qu, Jay Wang, Ling Zheng

机构 * Monmouth University(蒙茅斯大学)

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本文提出多模态网球运动员准备度预测框架PART,整合多源数据,可评估网球运动员健康、损伤风险等,在大学生网球运动员数据上表现良好,也适用于业余选手。

Comments 7 pages, 2 figures, 5 tables. Published in the 2025 IEEE 5th International Conference on Human-Machine Systems (ICHMS)

Journal ref 2025 IEEE 5th International Conference on Human-Machine Systems (ICHMS), pp. 28-34, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24903 2026-08-27 cs.HC cs.LG 新提交 78%

Evidence-Grounded Mapping of Multimodal Human Sensing Psychological Transdiagnostic Dimensions

基于证据的多模态人体感知心理跨诊断维度映射

Xiyun Hu, Xiangyuan Xue, Yuting Lyu, Hanya Shao, Jingping Nie

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本研究提出临床医生参与的基准,评估LLMs能否从多模态证据生成B-HiTOP条目画像,发现两阶段预测可提升部分证据的兼容性,但语义抽象会成为间接行为传感的信息瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21530 2026-08-25 cs.LG 新提交 78%

Multimodal Injury Risk and Performance Prediction in Tennis Using Weighted Ensemble Learning

基于加权集成学习的网球多模态损伤风险与表现预测

Weihao Qu, Dongyang Wang, Ling Zheng, Francisco E. Alvarez, Shobharani Polasa, Jiacun Wang

机构 * Monmouth University(蒙茅斯大学)

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本研究针对网球领域多模态损伤风险与表现预测方法不足的问题,提出PART多模态加权集成学习框架,整合多类数据提取专属特征,采用自适应权重策略,经9名大学网球运动员数据验证,可监测健康、估算损伤风险,对休闲球员也有应用潜力。

Comments 8 pages. Accepted author manuscript. Published as Early Access in IEEE Systems, Man, and Cybernetics Magazine

Journal ref IEEE Systems, Man, and Cybernetics Magazine, Early Access, pp. 1-7, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏