arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-03 至 2025-12-03 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 6 篇

2512.02405 2025-12-03 cs.CV cs.AI cs.LG 84%

WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate

WISE: 加权迭代专家社会用于鲁棒多模态多智能体辩论

Anoop Cherian, River Doyle, Eyal Ben-Dov, Suhas Lohit, Kuan-Chuan Peng

机构 * Mitsubishi Electric Research Labs(三菱电机研究实验室) Cambridge Rindge and Latin School(剑桥林恩和拉丁学校)

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出WISE框架,通过多模态多智能体辩论提升视觉语言推理任务的准确性,实验显示在多个数据集上提升了2-7%的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02981 2025-12-03 cs.CV 83%

InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration

InEx:通过内省与跨模态多智能体协作缓解幻觉

Zhongyu Yang, Yingfang Yuan, Xuanming Jiang, Baoyi An, Wei Pang

专题命中 多模态Agent :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 InEx通过内省推理和跨模态多智能体协作,自主缓解大型语言模型的幻觉问题,实验表明其在多个基准上表现优异。

Comments Published in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02777 2025-12-03 cs.RO cs.MA 78%

CogDrive: Cognition-Driven Multimodal Prediction-Planning Fusion for Safe Autonomy

CogDrive: 基于认知的多模态预测-规划融合用于安全自主性

Heye Huang, Yibin Yang, Mingfeng Fan, Haoran Wang, Xiaocong Zhao, Jianqiang Wang

机构 * Singapore-MIT Alliance for Research and Technology (SMART), Singapore(新加坡-麻省理工联合研究技术联盟) Department of Urban Studies and Planning, Massachusetts Institute of Technology, USA(麻省理工学院城市研究与规划系) School of Vehicle and Mobility, Tsinghua University, China(清华大学车辆与移动系统学院) Department of Mechanical Engineering, National University of Singapore, Singapore(新加坡国立大学机械工程系) Key Laboratory of Road and Traffic Engineering, Ministry of Education, Tongji University, China(同济大学交通工程重点实验室)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 CogDrive通过结合认知多模态预测与安全导向规划,实现了在复杂交通中的安全自主性,提升了轨迹预测和适应性行为。

Comments 25 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02280 2025-12-03 cs.AI cs.CV 62%

Bridging the Gap: Toward Cognitive Autonomy in Artificial Intelligence

弥合差距:迈向人工智能的认知自主性

Noorbakhsh Amiri Golilarz, Sindhuja Penchala, Shahram Rahimi

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文探讨了人工智能在自我监控、自我纠正和自主调节方面的能力不足,并提出基于神经认知原理的架构以实现认知自主性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02533 2025-12-03 cs.MM 57%

PopSim: Social Network Simulation for Social Media Popularity Prediction

PopSim: 基于社交网络的社交媒体流行度预测仿真

Yijun Liu, Wu Liu, Xiaoyan Gu, Allen He, Weiping Wang, Yongdong Zhang

专题命中 多模态Agent :multimodal(abstract);分类 cs.MM

AI总结 PopSim通过基于大语言模型的多智能体仿真,有效模拟UGC传播动态,提升社交媒体流行度预测的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01167 2025-12-03 cs.LG cs.AI cs.SY eess.SY 57%

A TinyML Reinforcement Learning Approach for Energy-Efficient Light Control in Low-Cost Greenhouse Systems

为低成本温室系统设计一种 TinyML 强化学习方法以实现节能照明控制

Mohamed Abdallah Salem, Manuel Cuevas Perez, Ahmed Harb Rabia

机构 * North Dakota State University(北达科他州立大学) Biosystems Engineering(生物系统工程)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出了一种基于 TinyML 的强化学习方法,用于低成本温室系统的节能照明控制,通过 Q 学习算法实现动态亮度调节,有效稳定不同光照水平。

Comments Copyright 2025 IEEE. This is the author's version of the work that has been accepted for publication in Proceedings of the 5. Interdisciplinary Conference on Electrics and Computer (INTCEC 2025) 15-16 September 2025, Chicago-USA. The final version of record is available at: https://doi.org/10.1109/INTCEC65580.2025.11256135

详情

展开后加载摘要…

URL PDF HTML 收藏