arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-30 至 2026-01-30 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 8 篇

2502.20120 2026-01-30 cs.CV 83%

Rethinking Multimodal Learning from the Perspective of Mitigating Classification Ability Disproportion

从缓解分类能力失衡角度重新思考多模态学习

QingYuan Jiang, Longfei Huang, Yang Yang

机构 * Nanjing University of Science and Technology(南京理工大学) State Key Lab. for Novel Software Technology, Nanjing University(南京大学软件新技术国家重点实验室)

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种基于提升原理的多模态学习方法,通过动态平衡弱强模态的分类能力以缓解模态不平衡问题。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.07292 2026-01-30 math.OC 78%

Path-Based Formulations for the Design of On-demand Multimodal Transit Systems with Adoption Awareness

基于路径的on-demand多模式交通系统设计的 formulations 与采纳意识

Hongzhao Guan, Beste Basciftci, Pascal Van Hentenryck

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本文提出了一种基于路径的优化模型P-Path,用于解决on-demand多模式交通系统设计中的计算难题,显著提高了求解效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02820 2026-01-30 cs.CV cs.AI cs.GR 62%

Mesh Neural Cellular Automata

网格神经元细胞自动机

Ehsan Pajouheshgar, Yitao Xu, Alexander Mordvintsev, Eyvind Niklasson, Tong Zhang, Sabine Süsstrunk

机构 * EPFL(瑞士联邦理工学院) Google Research(谷歌研究)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 MeshNCA是一种无需UV映射即可实时生成高质量3D动态纹理的神经元细胞自动机方法,通过多模态监督和用户交互实现了纹理合成的泛化能力。

Comments ACM Transactions on Graphics (TOG) - SIGGRAPH 2024

Journal ref ACM Transactions on Graphics (TOG), Volume 43, Issue 4 Article No.: 122, Pages 1 - 16; 19 July 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06753 2026-01-30 cs.CL 57%

Towards Computational Chinese Paleography

迈向计算中文古文字学

Yiran Rex Ma

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

AI总结 本文探讨了人工智能如何推动中文古文字学从孤立视觉任务向集成数字生态系统发展,强调多模态、少样本和以人为本的系统设计以解决数据稀缺和人文研究需求的挑战。

Comments A position paper in progress with Peking University & ByteDance Digital Humanities Open Lab

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21454 2026-01-30 cs.RO cs.CV 57%

4D-CAAL: 4D Radar-Camera Calibration and Auto-Labeling for Autonomous Driving

4D-CAAL:面向自动驾驶的4D雷达-相机校准与自动标注

Shanliang Yao, Zhuoxiao Li, Runwei Guan, Kebin Cao, Meng Xia, Fuping Hu, Sen Xu, Yong Yue, Xiaohui Zhu, Weiping Ding, Ryan Wen Liu

机构 * School of Information Engineering, Yancheng Institute of Technology(信息工程学院,盐城职业技术学院) School of Navigation, Wuhan University of Technology(导航学院,武汉理工大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) School of Information Engineering, Yancheng Institute Technology(信息工程学院,盐城职业技术学院) School of Advanced Technology, Xi’an Jiaotong-Liverpool University(先进技术学院,西安交通大学利物浦大学) School of Information Science and Technology, Nantong University(信息科学与技术学院,南通大学) State Key Laboratory of Maritime Technology and Safety(船舶技术与安全国家重点实验室)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 4D-CAAL提出了一种统一框架,通过双用途校准目标和自动标注流程,实现4D雷达与相机的高精度校准,减少人工标注工作量,加速自动驾驶多模态感知系统开发。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21316 2026-01-30 cs.LG cs.AI 57%

Heterogeneous Vertiport Selection Optimization for On-Demand Air Taxi Services: A Deep Reinforcement Learning Approach

异构垂直起降点选择优化用于按需空中出租车服务:一种深度强化学习方法

Aoyu Pang, Maonan Wang, Zifan Sha, Wenwei Yue, Changle Li, Chung Shue Chen, Man-On Pun

机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)科学与工程学院) Shanghai AI Laboratory(上海人工智能实验室) State Key Laboratory of Integrated Services Networks, Xidian University(西安电子科技大学集成服务网络国家重点实验室) Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出一种基于深度强化学习的框架,用于优化空中出租车服务中的垂直起降点选择,提升城市多模式交通系统的效率和整合性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22075 2026-01-30 cs.NE 50%

Lens-descriptor guided evolutionary algorithm for optimization of complex optical systems with glass choice

基于透镜描述符的进化算法用于具有玻璃选择的复杂光学系统优化

Kirill Antonov, Teus Tukker, Tiago Botari, Thomas H. W. Bäck, Anna V. Kononova, Niki van Stein

专题命中 其他多模态 :multimodal(abstract)

AI总结 提出LDG-EA算法,通过行为描述符和概率模型优化复杂光学系统,生成多样化的高质量透镜设计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21001 2026-01-30 cs.HC 50%

Designing the Interactive Memory Archive (IMA): A Socio-Technical Framework for AI-Mediated Reminiscence and Cultural Memory Preservation

设计交互记忆档案馆(IMA):一种面向AI辅助回忆和文化记忆保存的社会技术框架

Ron Fulbright

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文提出了一种基于AI的交互记忆档案馆框架,通过多模态感知和对话支架技术,帮助记忆丧失老年人参与回忆并保存文化记忆。

Comments 10 pages, 2 figures, 45 references cited

详情

展开后加载摘要…

URL PDF HTML 收藏