arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4895 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4895 篇

2511.06266 2025-11-18 cs.CV 57%

Spatially-Aware Mixture of Experts with Log-Logistic Survival Modeling for Whole-Slide Images

Ardhendu Sekhar, Vasu Soni, Keshav Aske, Shivam Madnoorkar, Pranav Jeevan, Amit Sethi

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05955 2025-11-17 cs.CV cs.LG 57%

CSGaze: Context-aware Social Gaze Prediction

Surbhi Madan, Shreya Ghosh, Ramanathan Subramanian, Abhinav Dhall, Tom Gedeon

机构 * Indian Institute of Technology Ropar(印度理工学院罗帕尔分校) The University of Queensland(昆士兰大学) University of Canberra(堪培拉大学) Curtin University(Curtin大学) Monash University(莫纳什大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10914 2025-11-17 cs.CV 57%

PhaseWin Search Framework Enable Efficient Object-Level Interpretation

Zihan Gu, Ruoyu Chen, Junchi Zhang, Yue Hu, Hua Zhang, Xiaochun Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) School of Mathematical Sciences, Fudan University(复旦大学数学学院) School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络科学与技术学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10890 2025-11-17 cs.AI stat.ML 57%

LLM enhanced graph inference for long-term disease progression modelling

Tiantian He, An Zhao, Elinor Thompson, Anna Schroder, Ahmed Abdulaal, Frederik Barkhof, Daniel C. Alexander

机构 * Department of Computer Science(计算机科学系)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10648 2025-11-14 cs.CV 57%

Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling

Jiahao Wang, Weiye Xu, Aijun Yang, Wengang Zhou, Lewei Lu, Houqiang Li, Xiaohua Wang, Jinguo Zhu

机构 * Xi’an Jiaotong University(西安交通大学) University of Science and Technology of China(中国科学技术大学) SenseTime Research(商汤科技研究院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025 (The Thirty-Ninth Annual Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09944 2025-11-14 cs.CV 57%

TSPE-GS: Probabilistic Depth Extraction for Semi-Transparent Surface Reconstruction via 3D Gaussian Splatting

Zhiyuan Xu, Nan Min, Yuhang Guo, Tong Wei

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments AAAI26 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05162 2025-11-14 cs.CY cs.AI 57%

Artificial-Intelligence Grading Assistance for Handwritten Components of a Calculus Exam

Gerd Kortemeyer, Alexander Caspar, Daria Horica

机构 * Michigan State University(密歇根州立大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07912 2025-11-12 cs.AI 57%

Neurophysiological Characteristics of Adaptive Reasoning for Creative Problem-Solving Strategy

Jun-Young Kim, Young-Seok Kweon, Gi-Hwan Shin, Seong-Whan Lee

机构 * Dept. of Artificial Intelligence(人工智能系) Korea University(韩国大学) Dept. of Brain and Cognitive Engineering(脑科学与认知工程系)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 4 pages, 4 figures, 1 table,

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18108 2025-11-12 cs.CV 57%

Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach

Jing Bi, Junjia Guo, Yunlong Tang, Lianggong Bruce Wen, Zhang Liu, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) Corning Inc(康宁公司)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Journal ref CVPR 2025 (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07004 2025-11-11 cs.CV cs.HC 57%

Exploring the "Great Unseen" in Medieval Manuscripts: Instance-Level Labeling of Legacy Image Collections with Zero-Shot Models

Christofer Meinecke, Estelle Guéville, David Joseph Wrisley

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06744 2025-11-11 cs.CV 57%

PointCubeNet: 3D Part-level Reasoning with 3x3x3 Point Cloud Blocks

Da-Yeong Kim, Yeong-Jun Cho

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06297 2025-11-11 cs.HC cs.AI 57%

Decomate: Leveraging Generative Models for Co-Creative SVG Animation

Jihyeon Park, Jiyoon Myung, Seone Shin, Jungki Son, Joohyung Han

机构 * MODULABS

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted at the 1st Workshop on Generative and Protective AI for Content Creation (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06256 2025-11-11 cs.CV 57%

VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving

Ruifei Zhang, Wei Zhang, Xiao Tan, Sibei Yang, Xiang Wan, Xiaonan Luo, Guanbin Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院) Sun Yat-sen University(中山大学) Baidu Inc.(百度公司) Guilin University of Electronic Technology(桂林电子科技大学) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)

专题命中 其他多模态 :MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13765 2025-11-06 cs.HC cs.CL 57%

SciDaSynth: Interactive Structured Data Extraction from Scientific Literature with Large Language Model

Xingbo Wang, Samantha L. Huey, Rui Sheng, Saurabh Mehta, Fei Wang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

Comments Preprint version of the paper accepted to Campbell Systematic Reviews. Code is available at https://github.com/xingbow/SciDaEx

Journal ref Campbell Systematic Reviews 21 (2025): 1-16

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09163 2025-11-05 cs.CV 57%

CWSSNet: Hyperspectral Image Classification Enhanced by Wavelet Domain Convolution

Yulin Tong, Fengzong Zhang, Haiqin Cheng

机构 * School of Transportation Engineering(交通运输工程学院) East China Jiaotong University(东华交通大学) School of Coumputing and Information(计算与信息学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.14269 2025-10-31 cs.LG cs.CV cs.IR 57%

Deep Learning for Technical Document Classification

Shuo Jiang, Jie Hu, Christopher L. Magee, Jianxi Luo

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments 16 pages, 8 figures, 9 tables

Journal ref IEEE Transactions on Engineering Management 71 (2024): 1163-1179

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10940 2025-10-30 cs.IR cs.AI 57%

Who You Are Matters: Bridging Topics and Social Roles via LLM-Enhanced Logical Recommendation

Qing Yu, Xiaobei Wang, Shuchang Liu, Yandong Bai, Xiaoyu Yang, Xueliang Wang, Chang Meng, Shanshan Wu, Hailan Yang, Huihui Xiao, Xiang Li, Fan Yang, Xiaoqiang Feng, Lantao Hu, Han Li, Kun Gai, Lixin Zou

机构 * Wuhan University(武汉大学) Kuaishou Technology(快手科技)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments to be published in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24023 2025-10-29 cs.CL 57%

Success and Cost Elicit Convention Formation for Efficient Communication

Saujas Vaduguru, Yilun Hua, Yoav Artzi, Daniel Fried

机构 * Carnegie Mellon University(卡内基梅隆大学) Department of Computer Science and Cornell Tech, Cornell University(计算机科学系和康奈尔科技,康奈尔大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23648 2025-10-29 cs.SI cs.AI 57%

RoGBot: Relationship-Oblivious Graph-based Neural Network with Contextual Knowledge for Bot Detection

Ashutosh Anshul, Mohammad Zia Ur Rehman, Sri Akash Kadali, Nagendra Kumar

机构 * Indian Institute of Technology Indore(印度理工学院印多尔分校) University of Maryland, College Park, USA(美国马里兰大学学院公园分校)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Submitted to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17047 2025-10-28 cs.CL 57%

Modeling Bottom-up Information Quality during Language Processing

Cui Ding, Yanning Yin, Lena A. Jäger, Ethan Gotlieb Wilcox

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02366 2025-10-28 cs.LG cs.CL q-fin.TR 57%

Language Model Guided Reinforcement Learning in Quantitative Trading

Adam Darmanin, Vince Vella

机构 * University of Malta(马耳他大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL

Comments 12 pages (4 pages appendix and references) and 6 figures. Accepted for presentation at FLLM 2025, Vienna

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14677 2025-10-28 cs.CV 57%

Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning

Jiaer Xia, Yuhang Zang, Peng Gao, Sharon Li, Kaiyang Zhou

机构 * Hong Kong Baptist University(香港 Baptist 大学) Shanghai AI Lab(上海人工智能实验室) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17611 2025-10-27 cs.CV 57%

One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection

Jia Guo, Shuai Lu, Lei Fan, Zelin Li, Donglin Di, Yang Song, Weihang Zhang, Wenbing Zhu, Hong Yan, Fang Chen, Huiqi Li, Hongen Liao

机构 * Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学) Shanghai Jiao Tong University(上海交通大学) City University of Hong Kong(香港城市大学) University of New South Wales(新南威尔士大学) DZ Matrix(DZ矩阵) Fudan University(复旦大学) Rongcheer Co., Ltd.(荣彻科技有限公司)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments Extended version of CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16034 2025-10-21 cs.CV 57%

VisualLens: Personalization through Task-Agnostic Visual History

Wang Bill Zhu, Deqing Fu, Kai Sun, Yi Lu, Zhaojiang Lin, Seungwhan Moon, Kanika Narang, Mustafa Canim, Yue Liu, Anuj Kumar, Xin Luna Dong

机构 * Meta University of Southern California(南加州大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13660 2025-10-17 cs.CV 57%

OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild

Hongyu Qu, Jianan Wei, Xiangbo Shu, Yazhou Yao, Wenguan Wang, Jinhui Tang

机构 * Nanjing University of Science and Technology(南京理工大学) Zhejiang University(浙江大学) Nanjing Forestry University(南京林业大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025; Project page: https://github.com/quhongyu/OmniGaze

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00527 2025-10-17 cs.RO cs.AI 57%

Never too Prim to Swim: An LLM-Enhanced RL-based Adaptive S-Surface Controller for AUVs under Extreme Sea Conditions

Guanwen Xie, Jingzehua Xu, Yimian Ding, Zhi Zhang, Shuai Zhang, Yi Li

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, 518055, China(清华大学深圳国际研究生院,清华大学,深圳,518055,中国) Department of Data Science, New Jersey Institute of Technology(数据科学系,新泽西理工学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted by IEEE/RSJ IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13029 2025-10-16 cs.AI 57%

Toward Reasoning-Centric Time-Series Analysis

Xinlei Wang, Mingtian Tan, Jing Qiu, Junhua Zhao, Jinjin Gu

机构 * University of Sydney(悉尼大学) University of Virginia(弗吉尼亚大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Institute of Artificial Intelligence and Robotics for Society(深圳人工智能与机器人社会研究院) INSAIT

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23179 2025-10-16 cs.CV 57%

DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes

Sungjune Park, Hyunjun Kim, Junho Kim, Seongho Kim, Yong Man Ro

机构 * Integrated Vision and Language Lab., School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(整合视觉与语言实验室,电气工程学院,韩国科学技术院(KAIST))

专题命中 其他多模态 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11106 2025-10-14 cs.CV 57%

Compositional Zero-Shot Learning: A Survey

Ans Munir, Faisal Z. Qureshi, Mohsen Ali, Muhammad Haris Khan

机构 * Information Technology University(信息科技大学) University of Ontario Institute of Technology(安大略理工大学) Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

Comments Survey paper with 36 pages, 8 plots and 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21988 2025-10-09 cs.AI 57%

Functional Matching of Logic Subgraphs: Beyond Structural Isomorphism

Ziyang Zheng, Kezhi Li, Zhengyuan Shi, Qiang Xu

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏