arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-01 至 2025-12-01 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 6 篇

2511.21902 2025-12-01 cs.CV cs.AI 81%

PathReasoning: A multimodal reasoning agent for query-based ROI navigation on whole-slide images

PathReasoning: 一种用于基于查询的全滑片图像区域感兴趣点(ROI)导航的多模态推理代理

Kunpeng Zhang, Hanwen Xu, Sheng Wang

专题命中 多模态Agent :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 PathReasoning通过多模态推理代理实现基于查询的全滑片图像ROI导航,显著提升诊断准确性与报告生成效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22181 2025-12-01 cs.CV cs.AI cs.RO 62%

MTR-VP: Towards End-to-End Trajectory Planning through Context-Driven Image Encoding and Multiple Trajectory Prediction

MTR-VP: 通过基于上下文的图像编码和多轨迹预测实现端到端轨迹规划

Maitrayee Keskar, Mohan Trivedi, Ross Greer

机构 * Machine Intelligence, Interaction, and Imagination (Mi 3 ) Laboratory(机器智能、交互与想象实验室) University of California, Merced(加州大学默塞德分校) Laboratory for Intelligent & Safe Automobiles (LISA)(智能与安全汽车实验室) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 MTR-VP通过基于上下文的图像编码和多轨迹预测实现端到端轨迹规划,利用交叉注意力提升规划性能。

Comments 8 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22873 2025-12-01 cs.CV cs.IR 57%

CNN-Based Framework for Pedestrian Age and Gender Classification Using Far-View Surveillance in Mixed-Traffic Intersections

基于CNN的远视监控中混合交通交叉口行人年龄和性别分类框架

Shisir Shahriar Arif, Md. Muhtashim Shahrier, Nazmul Haque, Md Asif Raihan, Md. Hadiuzzaman

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 本研究提出基于CNN的远视监控框架,用于混合交通交叉口行人年龄和性别分类,无需面部识别或高分辨率图像,提供高效、低成本的行人人口统计数据监测解决方案。

Comments Accepted for poster presentation at the 105th Annual Meeting of the Transportation Research Board

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22737 2025-12-01 cs.AI cs.HC 57%

Agentic AI Framework for Individuals with Disabilities and Neurodivergence: A Multi-Agent System for Healthy Eating, Daily Routines, and Inclusive Well-Being

具有残疾和神经多样性个体的代理AI框架:一个用于健康饮食、日常习惯和包容性福祉的多代理系统

Salman Jan, Toqeer Ali Syed, Gohar Ali, Ali Akarma, Mohammad Riyaz Belgaum, Ahmad Ali

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出一种多代理系统,通过个性化营养、适应性调度、食品指导和生理监测等代理,为残疾和神经多样性个体提供健康饮食、日常习惯和包容性福祉的AI框架。

Comments Presented at International Conference on Business and Digital Technology, Bahrain, Springer Nature, 27 November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22134 2025-12-01 cs.CV cs.RO 57%

DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action

DualVLA: 通过推理与行动部分解耦构建通用具身代理

Zhen Fang, Zhuoyang Liu, Jiaming Liu, Hao Chen, Yu Zeng, Shiting Huang, Zehui Chen, Lin Chen, Shanghang Zhang, Feng Zhao

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知国家重点实验室,中国科学技术大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,北京大学计算机学院) CUHK(香港大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 DualVLA通过推理与行动部分解耦,提升通用具身代理的行动与多模态理解平衡能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22390 2025-12-01 cs.LO 50%

Modal Logic for Simulation, Refinement, and Mutual Ignorance

模态逻辑用于模拟、细化和相互无知

Hans van Ditmarsch, Tim French, Rustam Galimullin, Louwe B. Kuijer

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出了一种基于多模态逻辑的模态逻辑,用于模拟、细化和相互无知,通过模块化的方式构建了多种逻辑体系,探讨了细化与模拟之间的关系。

Comments In Proceedings TARK 2025, arXiv:2511.20540

Journal ref EPTCS 437, 2025, pp. 379-398

详情

展开后加载摘要…

URL PDF HTML 收藏