arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-01 至 2025-12-01 共收录 111 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 20 篇

2510.23240 2025-12-01 cs.CV 57%

Autoregressive Styled Text Image Generation, but Make it Reliable

自回归风格文本图像生成,但使其更可靠

Carmine Zaccagnino, Fabio Quattrini, Vittorio Pippi, Silvia Cascianelli, Alessio Tonioni, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) Google(谷歌)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出Eruku方法,通过多模态提示条件生成任务和分类器自由引导策略,改进自回归模型以生成更可靠、更忠实于文本提示的风格化文本图像。

Comments Accepted at WACV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00802 2025-12-01 cs.CV 57%

TRACE: Temporally Reliable Anatomically-Conditioned 3D CT Generation with Enhanced Efficiency

TRACE: 基于时间可靠性的解剖条件3D CT生成与增强效率

Minye Shao, Xingyu Miao, Haoran Duan, Zeyu Wang, Jingkun Chen, Yawen Huang, Xian Wu, Jingjing Deng, Yang Long, Yefeng Zheng

机构 * Department of Computer Science, Durham University(杜伦大学计算机科学系) Department of Automation, Tsinghua University(清华大学自动化系) College of Computer Science and Engineering, Dalian Minzu University(大连民族大学计算机科学与工程学院) Department of Engineering Science, University of Oxford(牛津大学工程科学系) Jarvis Research Center, Tencent YouTu Lab(腾讯YouTu实验室 Jarvis 研究中心) School of Engineering Mathematics and Technology, University of Bristol(布里斯托大学工程数学与技术学院) Medical Artificial Intelligence Laboratory, School of Engineering, Westlake University(西湖大学工程学院医学人工智能实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 TRACE通过2D多模态条件扩散方法生成具有时空对齐的3D CT图像,提升生成效率和解剖保真度。

Comments Accepted to MICCAI 2025 (this version is not peer-reviewed; it is the extended version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21780 2025-12-01 cs.MM cs.SD 57%

3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation

3MDiT:用于文本驱动同步音频视频生成的统一三模态扩散变换器

Yaoru Li, Heyu Si, Federico Landi, Pilar Oplustil Gallegos, Ioannis Koutsoumpas, O. Ricardo Cortez Vazquez, Ruiju Fu, Qi Guo, Xin Jin, Shunyu Liu, Mingli Song

机构 * Zhejiang University(浙江大学) Huawei(华为) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.MM

AI总结 3MDiT提出了一种统一三模态扩散变换器,通过联合演变流实现文本驱动的同步音频视频生成,提升多模态对齐与同步性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21746 2025-12-01 cs.CL 57%

DELTA: Language Diffusion-based EEG-to-Text Architecture

DELTA: 基于语言扩散的EEG到文本架构

Mingyu Jeon, Hyobin Kim

机构 * Sungkyunkwan University(成均馆大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

AI总结 DELTA通过结合残差向量量化和语言扩散模型,提升了EEG到文本的语义对齐和生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02058 2025-12-01 q-bio.BM cs.LG 50%

RiboGen: RNA Sequence and Structure Co-Generation with Equivariant MultiFlow

RiboGen:基于等变多流的RNA序列和结构联合生成

Dana Rubin, Allan dos Santos Costa, Manvitha Ponnapati, Joseph Jacobson

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) MIT Media Lab(麻省理工学院媒体实验室) Molecular Machines(分子机器) Center for Bits and Atoms(比特与原子中心)

专题命中 多模态生成 :multimodal(abstract)

AI总结 RiboGen通过等变多流方法实现RNA序列和结构的联合生成,为RNA设计提供了新的深度学习解决方案。

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态评测 18 篇

2511.22055 2025-12-01 cs.CV cs.MM 84%

OralGPT-Omni: A Versatile Dental Multimodal Large Language Model

OralGPT-Omni: 一种多功能的牙科多模态大语言模型

Jing Hao, Yuci Liang, Lizhuo Lin, Yuxuan Fan, Wenkai Zhou, Kaixin Guo, Zanting Ye, Yanpeng Sun, Xinyu Zhang, Yanqi Yang, Qiankun Li, Hao Tang, James Kit-Hon Tsoi, Linlin Shen, Kuo Feng Hung

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院) College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学) School of Biomedical Engineering, Southern Medical University(南方医科大学生物医学工程学院) Singapore University of Technology and Design(新加坡科技与设计大学) University of Auckland(奥克兰大学) University of Science and Technology of China(中国科学技术大学) School of Computer Science, Peking University(北京大学计算机学院) College of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.MM

AI总结 OralGPT-Omni是一种专门用于牙科的多模态大语言模型,通过TRACE-CoT数据集和四阶段训练范式,实现了对牙科图像的高效分析和高准确率评估。

Comments 47 pages, 42 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22404 2025-12-01 cs.CV 83%

UAV-MM3D: A Large-Scale Synthetic Benchmark for 3D Perception of Unmanned Aerial Vehicles with Multi-Modal Data

UAV-MM3D: 一种大规模合成基准,用于多模态数据下的无人机三维感知

Longkun Zou, Jiale Wang, Rongqin Liang, Hai Wu, Ke Chen, Yaowei Wang

机构 * Pengcheng Laboratory(鹏城实验室) University of Southern California(南加州大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 UAV-MM3D通过多模态合成数据提升无人机三维感知能力,提供高保真数据集和多任务基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07984 2025-12-01 cs.CV 83%

SAMChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Small Scale Remote Sensing

SAMChat:引入链式推理和GRPO以增强小规模遥感遥感小语言模型

Aybora Koksal, A. Aydin Alatan

机构 * Center for the Image Analysis (OGAM) and Department of Electrical and Electronics Engineering of Middle East Technical University (METU)(图像分析中心(OGAM)和中东部技术大学(METU)电子与电气工程系)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 SAMChat通过引入链式推理和GRPO,专为遥感影像分析优化,实现了在开放描述和分类任务上的高精度表现。

Comments Accepted to Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS) Special Issue on Foundation and Large Vision Models for Remote Sensing. Code and dataset are available at https://github.com/aybora/SAMChat

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12861 2025-12-01 cs.CL cs.CV 81%

From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models

从感知到推理:深度思考赋能多模态大语言模型

Wenxin Zhu, Andong Chen, Yuchen Song, Kehai Chen, Conghui Zhu, Ziyan Chen, Tiejun Zhao

机构 * Harbin Institute of Technology(哈尔滨工业大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Global Tone Communication Technology(全球语音通信技术)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出多模态链式思维方法,旨在提升多模态大语言模型的推理能力,通过系统性综述分析其理论基础、实现方法及未来发展方向。

Comments Survey; 7 figures, 3 tables, 44 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23269 2025-12-01 cs.AI 79%

OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

OctoMed:面向尖端多模态医疗推理的数据配方

Timothy Ossowski, Sheng Zhang, Qianchu Liu, Guanghui Qin, Reuben Tan, Tristan Naumann, Junjie Hu, Hoifung Poon

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Microsoft Research(微软研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 OctoMed通过结构化推理轨迹的数据配方,提升医疗多模态推理模型的性能和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08303 2025-12-01 cs.CV 79%

Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers

通过多模态视觉序列变压器推进语义未来预测

Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis

机构 * Archimedes, Athena Research Center(阿基米德研究中心) National Technical University of Athens(希腊国家技术大学) University of Crete(克里特大学) IACM-Forth(第四研究机构(IACM-Forth))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 FUTURIST通过多模态视觉序列变压器架构实现高效的多模态未来语义预测,提升预测精度并简化训练流程。

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11414 2025-12-01 cs.CL 79%

Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization

细粒度且可解释的事实性评估用于多模态摘要

Yue Zhang, Jingxuan Zuo, Ke Su, Liqiang Jing

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出两种细粒度且可解释的评估框架,用于评估多模态摘要模型的事实性,适用于不同应用场景,并通过实验验证了其有效性。

Comments project link: https://github.com/for4WARD/FaithfulnessEvaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21795 2025-12-01 cs.CR 78%

Advanced Data Collection Techniques in Cloud Security: A Multi-Modal Deep Learning Autoencoder Approach

云安全中的高级数据收集技术:一种多模态深度学习自编码器方法

Aamiruddin Syed, Mohammed Ilyas Ahmad

专题命中 多模态评测 :multi-modal(title,abstract)

AI总结 本文提出一种多模态深度学习自编码器方法,通过整合多种数据源和模态,提升云安全中的异常检测与分类性能。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00369 2025-12-01 cs.CV 74%

Automated segmentation of pediatric neuroblastoma on multi-modal MRI: Results of the SPPIN challenge at MICCAI 2023

多模态MRI上儿科神经母细胞瘤自动分割:2023年MICCAI SPPIN挑战赛结果

M. A. D. Buser, D. C. Simons, M. Fitski, M. H. W. A. Wijnen, A. S. Littooij, A. H. ter Brugge, I. N. Vos, M. H. A. Janse, M. de Boer, R. ter Maat, J. Sato, S. Kido, S. Kondo, S. Kasai, M. Wodzinski, H. Muller, J. Ye, J. He, Y. Kirchhoff, M. R. Rokkus, G. Haokai, S. Zitong, M. Fernández Patón, D. Veiga-Canuto, D. G. Ellis, M. R. Aizenberg, B. H. M. van der Velden, H. Kuijf, A. De Luca, A. F. W. van der Steeg

专题命中 多模态评测 :multi-modal(title);分类 cs.CV

AI总结 SPPIN挑战赛通过多模态MRI自动分割神经母细胞瘤,展示了预训练网络在小数据集中的有效性,但小肿瘤分割仍需改进。

Comments 23 pages, 6 figures

Journal ref Bioengineering, 12(11), 1157 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18842 2025-12-01 cs.CV 74%

Enhancing Descriptive Image Quality Assessment with A Large-scale Multi-modal Dataset

通过大规模多模态数据集增强描述性图像质量评估

Zhiyuan You, Jinjin Gu, Xin Cai, Zheyuan Li, Kaiwen Zhu, Chao Dong, Tianfan Xue

机构 * The Chinese University of Hong Kong(香港中文大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) Sofia University(索菲亚大学) University of Macau(澳门大学) Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 多模态评测 :multi-modal(title);分类 cs.CV

AI总结 本研究提出DepictQA-Wild模型,通过构建大规模多模态数据集提升图像质量评估的准确性和实用性。

Comments Accepted by TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11561 2025-12-01 cs.CV 70%

Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution

利用大规模语言模型回归准确的图像质量评分使用分数分布

Zhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue, Chao Dong

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Multimedia Laboratory, The Chinese University of Hong Kong(香港中文大学多媒体实验室) Shanghai AI Laboratory(上海人工智能实验室) Shenzhen University of Advanced Technology(深圳先进技术大学) CPII under InnoHK(创新香港下的CPII)

专题命中 多模态评测 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本研究提出基于分布的DeQA-Score模型,通过离散化评分分布为软标签,提升图像质量评分的准确性和一致性。

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21735 2025-12-01 cs.CL cs.AI cs.CV 67%

Closing the Performance Gap Between AI and Radiologists in Chest X-Ray Reporting

弥合AI与放射科医生在胸部X光报告中的性能差距

Harshita Sharma, Maxwell C. Reynolds, Valentina Salvatelli, Anne-Marie G. Sykes, Kelly K. Horst, Anton Schwaighofer, Maximilian Ilse, Olesya Melnichenko, Sam Bond-Taylor, Fernando Pérez-García, Vamshi K. Mugu, Alex Chan, Ceylan Colak, Shelby A. Swartz, Motassem B. Nashawaty, Austin J. Gonzalez, Heather A. Ouellette, Selnur B. Erdal, Beth A. Schueler, Maria T. Wetscherek, Noel Codella, Mohit Jain, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Stephanie Hyland, Panos Korfiatis, Ashish Khandelwal, Javier Alvarez-Valle

机构 * Microsoft(微软公司) Mayo Clinic(梅奥诊所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 MAIRA-X通过多模态AI模型在胸部X光报告生成中提升词汇质量、临床正确性和L&T准确性,有效辅助放射科医生,尤其在高患者量的临床环境中

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22178 2025-12-01 cs.CV cs.AI 62%

Enhanced Graph Convolutional Network with Chebyshev Spectral Graph and Graph Attention for Autism Spectrum Disorder Classification

增强型图卷积网络结合切比雪夫谱图与图注意力用于自闭症谱系障碍分类

Adnan Ferdous Ashrafi, Hasanul Kabir

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Islamic University of Technology(伊斯兰科技大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出结合切比雪夫谱图卷积和图注意力网络的增强型图卷积网络,用于提高自闭症谱系障碍分类的准确性。

Comments 6 pages, 2 figures, 2 tables, Accepted and presented at Image and Vision Computing New Zealand (IVCNZ) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23214 2025-12-01 cs.CV 57%

Zero-Shot Multi-Criteria Visual Quality Inspection for Semi-Controlled Industrial Environments via Real-Time 3D Digital Twin Simulation

面向半受控工业环境的零样本多准则视觉质量检测:通过实时3D数字孪生模拟

Jose Moises Araya-Martinez, Gautham Mohan, Kenichi Hayakawa Bolaños, Roberto Mendieta, Sarvenaz Sardari, Jens Lambrecht, Jörg Krüger

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种基于实时3D数字孪生模拟的零样本多准则视觉质量检测框架,用于半受控工业环境中的高效质量检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22576 2025-12-01 cs.MM 57%

A Progressive Evaluation Framework for Multicultural Analysis of Story Visualization

一种用于故事可视化多元文化分析的渐进评估框架

Janak Kapuriya, Ali Hatami, Paul Buitelaar

专题命中 多模态评测 :MLLM(abstract);分类 cs.MM

AI总结 本文提出一种渐进多元文化评估框架,通过五个新指标评估故事可视化模型在不同文化下的表现,揭示模型在文化适当性和视觉美学上的差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21716 2025-12-01 cs.CL 57%

An Optimized Machine Learning Classifier for Detecting Fake Reviews Using Extracted Features

一种用于通过提取特征检测虚假评论的优化机器学习分类器

Shabbir Anees, Anshuman, Ayush Chaurasia, Prathmesh Bogar

机构 * Indian Institute of Information Technology Vadodara(印度瓦达拉信息科技大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL

AI总结 本文提出了一种结合HHO优化和堆叠集成的机器学习方法,用于高精度识别由人工智能生成的虚假评论。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06597 2025-12-01 cs.RO 50%

LiHRA: A LiDAR-Based HRI Dataset for Automated Risk Monitoring Methods

LiHRA:基于LiDAR的人机交互风险监测数据集

Frederik Plahl, Georgios Katranis, Ilshat Mamaev, Andrey Morozov

机构 * Proximity Robotics & Automation GmbH(近距机器人与自动化有限公司) Institute of Industrial Automation and Software Engineering, University of Stuttgart(工业自动化与软件工程学院,斯图加特大学)

专题命中 多模态评测 :multi-modal(abstract)

AI总结 LiHRA数据集通过多模态数据支持人机交互风险监测方法的开发,提供高分辨率LiDAR数据和真实碰撞事件,用于训练和评估RM算法。

Comments Preprint of final paper that will appear in the Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09042 2025-12-01 eess.SY cs.LG cs.SY 50%

MAKO: Meta-Adaptive Koopman Operators for Learning-based Model Predictive Control of Parametrically Uncertain Nonlinear Systems

MAKO:基于元学习的Koopman算子用于参数不确定非线性系统的基于学习的模型预测控制

Minghao Han, Kiwan Wong, Adrian Wing-Keung Law, Xunyuan Yin

机构 * Water Research Institute (NEWRI), Nanyang Technological University, Singapore(新跃大学水研究 institute(NEWRI)) School of Chemistry, Chemical Engineering and Biotechnology, Nanyang Technological University, Singapore(新跃大学化学、化工与生物技术学院) Soft Robotics Lab, ETH Zurich, Switzerland(苏黎世联邦理工学院软机器人实验室) Department of Civil and Environmental Engineering, National University of Singapore, Singapore(新加坡国立大学土木与环境工程系)

专题命中 多模态评测 :multi-modal(abstract)

AI总结 MAKO通过元学习方法实现对参数不确定非线性系统的高效建模与预测控制,提升系统稳定性和控制效果。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态Agent 6 篇

2511.21902 2025-12-01 cs.CV cs.AI 81%

PathReasoning: A multimodal reasoning agent for query-based ROI navigation on whole-slide images

PathReasoning: 一种用于基于查询的全滑片图像区域感兴趣点(ROI)导航的多模态推理代理

Kunpeng Zhang, Hanwen Xu, Sheng Wang

专题命中 多模态Agent :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 PathReasoning通过多模态推理代理实现基于查询的全滑片图像ROI导航,显著提升诊断准确性与报告生成效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22181 2025-12-01 cs.CV cs.AI cs.RO 62%

MTR-VP: Towards End-to-End Trajectory Planning through Context-Driven Image Encoding and Multiple Trajectory Prediction

MTR-VP: 通过基于上下文的图像编码和多轨迹预测实现端到端轨迹规划

Maitrayee Keskar, Mohan Trivedi, Ross Greer

机构 * Machine Intelligence, Interaction, and Imagination (Mi 3 ) Laboratory(机器智能、交互与想象实验室) University of California, Merced(加州大学默塞德分校) Laboratory for Intelligent & Safe Automobiles (LISA)(智能与安全汽车实验室) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 MTR-VP通过基于上下文的图像编码和多轨迹预测实现端到端轨迹规划,利用交叉注意力提升规划性能。

Comments 8 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22873 2025-12-01 cs.CV cs.IR 57%

CNN-Based Framework for Pedestrian Age and Gender Classification Using Far-View Surveillance in Mixed-Traffic Intersections

基于CNN的远视监控中混合交通交叉口行人年龄和性别分类框架

Shisir Shahriar Arif, Md. Muhtashim Shahrier, Nazmul Haque, Md Asif Raihan, Md. Hadiuzzaman

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 本研究提出基于CNN的远视监控框架,用于混合交通交叉口行人年龄和性别分类,无需面部识别或高分辨率图像,提供高效、低成本的行人人口统计数据监测解决方案。

Comments Accepted for poster presentation at the 105th Annual Meeting of the Transportation Research Board

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22737 2025-12-01 cs.AI cs.HC 57%

Agentic AI Framework for Individuals with Disabilities and Neurodivergence: A Multi-Agent System for Healthy Eating, Daily Routines, and Inclusive Well-Being

具有残疾和神经多样性个体的代理AI框架:一个用于健康饮食、日常习惯和包容性福祉的多代理系统

Salman Jan, Toqeer Ali Syed, Gohar Ali, Ali Akarma, Mohammad Riyaz Belgaum, Ahmad Ali

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出一种多代理系统,通过个性化营养、适应性调度、食品指导和生理监测等代理,为残疾和神经多样性个体提供健康饮食、日常习惯和包容性福祉的AI框架。

Comments Presented at International Conference on Business and Digital Technology, Bahrain, Springer Nature, 27 November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22134 2025-12-01 cs.CV cs.RO 57%

DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action

DualVLA: 通过推理与行动部分解耦构建通用具身代理

Zhen Fang, Zhuoyang Liu, Jiaming Liu, Hao Chen, Yu Zeng, Shiting Huang, Zehui Chen, Lin Chen, Shanghang Zhang, Feng Zhao

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知国家重点实验室,中国科学技术大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,北京大学计算机学院) CUHK(香港大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 DualVLA通过推理与行动部分解耦,提升通用具身代理的行动与多模态理解平衡能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22390 2025-12-01 cs.LO 50%

Modal Logic for Simulation, Refinement, and Mutual Ignorance

模态逻辑用于模拟、细化和相互无知

Hans van Ditmarsch, Tim French, Rustam Galimullin, Louwe B. Kuijer

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出了一种基于多模态逻辑的模态逻辑,用于模拟、细化和相互无知,通过模块化的方式构建了多种逻辑体系,探讨了细化与模拟之间的关系。

Comments In Proceedings TARK 2025, arXiv:2511.20540

Journal ref EPTCS 437, 2025, pp. 379-398

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态训练与对齐 18 篇

2506.03195 2025-12-01 cs.CV cs.AI cs.LG 84%

Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs

未标记数据提升多模态大语言模型在细粒度图像零样本分类中的性能

Yunqi Hong, Sohyun An, Andrew Bai, Neil Y. C. Lin, Cho-Jui Hsieh

机构 * Computer Science Department, University of California, Los Angeles(加州大学洛杉矶分校计算机科学系) Mechanical and Aerospace Engineering Department, University of California, Los Angeles(加州大学洛杉矶分校机械与航空航天工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 AutoSEP通过利用未标记数据提升多模态大语言模型在细粒度图像零样本分类中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏