arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2407.10151 2024-12-09 cs.CV 57%

Lost and Found: Overcoming Detector Failures in Online Multi-Object Tracking

Lorenzo Vaquero, Yihong Xu, Xavier Alameda-Pineda, Victor M. Brea, Manuel Mucientes

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted at ECCV 2024. Code available at https://github.com/lorenzovaquero/BUSCA

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04069 2024-12-06 cs.AI 57%

ProtDAT: A Unified Framework for Protein Sequence Design from Any Protein Text Description

Xiao-Yu Guo, Yi-Fan Li, Yuan Liu, Xiaoyong Pan, Hong-Bin Shen

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02062 2024-12-04 cs.AI cs.CY 57%

Construction and optimization of health behavior prediction model for the elderly in smart elderly care

Qian Guo, Peiyuan Chen

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00651 2024-12-03 cs.CV q-bio.GN 57%

Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics

Minghao Han, Dingkang Yang, Jiabei Cheng, Xukun Zhang, Linhao Qu, Zizhi Chen, Lihua Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 21 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00447 2024-12-03 cs.CV 57%

ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models

Xubing Ye, Yukang Gan, Yixiao Ge, Xiao-Ping Zhang, Yansong Tang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02220 2024-12-02 cs.CV cs.CR cs.LG 57%

QuantAttack: Exploiting Dynamic Quantization to Attack Vision Transformers

Amit Baras, Alon Zolfi, Yuval Elovici, Asaf Shabtai

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17088 2024-11-27 cs.CV 57%

ΩSFormer: Dual-Modal Ω-like Super-Resolution Transformer Network for Cross-scale and High-accuracy Terraced Field Vectorization Extraction

Chang Li, Yu Wang, Ce Zhang, Yongjun Zhang

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16809 2024-11-27 cs.CR cs.AI 57%

Blockchain Meets LLMs: A Living Survey on Bidirectional Integration

Jianghao Gong, Peiqi Yan, Yue Zhang, Hongli An, Logan Liu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16561 2024-11-26 cs.SE cs.CL 57%

EnStack: An Ensemble Stacking Framework of Large Language Models for Enhanced Vulnerability Detection in Source Code

Shahriyar Zaman Ridoy, Md. Shazzad Hossain Shaon, Alfredo Cuzzocrea, Mst Shapna Akter

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CL

Comments Accepted in 2024 IEEE International Conference on Big Data (IEEE BigData 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15837 2024-11-26 cs.CV 57%

Modality Alignment Meets Federated Broadcasting

Yuting Ma, Shengeng Tang, Xiaohua Xu, Lechao Cheng

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15500 2024-11-26 cs.ET cs.CL 57%

MolMetaLM: a Physicochemical Knowledge-Guided Molecular Meta Language Model

Yifan Wu, Min Zeng, Yang Li, Yang Zhang, Min Li

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14768 2024-11-25 cs.LG cs.AI 57%

Grid and Road Expressions Are Complementary for Trajectory Representation Learning

Silin Zhou, Shuo Shang, Lisi Chen, Peng Han, Christian S. Jensen

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments This paper is accepted by KDD2025(August Cycle)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14228 2024-11-22 cs.CV 57%

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Yuke Zhu, Chi Xie, Shuang Liang, Bo Zheng, Sheng Guo

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.24001 2024-11-19 cs.CV 57%

ImOV3D: Learning Open-Vocabulary Point Clouds 3D Object Detection from Only 2D Images

Timing Yang, Yuanliang Ju, Li Yi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2024. Code link https://github.com/yangtiming/ImOV3D

Journal ref NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01412 2024-11-12 cs.CV cs.RO 57%

DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model

Zhenhua Xu, Yujia Zhang, Enze Xie, Zhen Zhao, Yong Guo, Kwan-Yee. K. Wong, Zhenguo Li, Hengshuang Zhao

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by RA-L. The project page is available at https://tonyxuqaq.github.io/projects/DriveGPT4/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05254 2024-11-11 cs.CV 57%

Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding

Jaeyoo Park, Jin Young Choi, Jeonghyung Park, Bohyung Han

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17367 2024-11-07 cs.LG cs.CV 57%

Implicit Neural Representations for Simultaneous Reduction and Continuous Reconstruction of Multi-Altitude Climate Data

Alif Bin Abdul Qayyum, Xihaier Luo, Nathan M. Urban, Xiaoning Qian, Byung-Jun Yoon

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:2401.16936

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17331 2024-11-07 cs.CV 57%

Multi-label Cluster Discrimination for Visual Representation Learning

Xiang An, Kaicheng Yang, Xiangzi Dai, Ziyong Feng, Jiankang Deng

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

Comments Accepted by ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02999 2024-11-06 cs.CV 57%

Precise Drive with VLM: First Prize Solution for PRCV 2024 Drive LM challenge

Bin Huang, Siyu Wang, Yuanpeng Chen, Yidan Wu, Hui Song, Zifan Ding, Jing Leng, Chengpeng Liang, Peng Xue, Junliang Zhang, Tiankun Zhao

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04802 2024-11-06 cs.CV cs.LG 57%

Predictive Dynamic Fusion

Bing Cao, Yinan Xia, Yi Ding, Changqing Zhang, Qinghua Hu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00290 2024-11-06 cs.CV 57%

MS-DETR: Multispectral Pedestrian Detection Transformer with Loosely Coupled Fusion and Modality-Balanced Optimization

Yinghui Xing, Shuo Yang, Song Wang, Shizhou Zhang, Guoqiang Liang, Xiuwei Zhang, Yanning Zhang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments The paper has been accepted by IEEE Transactions on Intelligent Transportation Systems

Journal ref IEEE Transactions on Intelligent Transportation Systems 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22982 2024-10-31 cs.RO cs.AI cs.SY eess.SY 57%

PDSR: Efficient UAV Deployment for Swift and Accurate Post-Disaster Search and Rescue

Alaa Awad Abdellatif, Ali Elmancy, Amr Mohamed, Ahmed Massoud, Wadha Lebda, Khalid K. Naji

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments This paper is currently under review at IEEE IoT Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09674 2024-10-31 eess.IV cs.CV cs.LG cs.NE 57%

EG-SpikeFormer: Eye-Gaze Guided Transformer on Spiking Neural Networks for Medical Image Analysis

Yi Pan, Hanqi Jiang, Junhao Chen, Yiwei Li, Huaqin Zhao, Yifan Zhou, Peng Shu, Zihao Wu, Zhengliang Liu, Dajiang Zhu, Xiang Li, Yohannes Abate, Tianming Liu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13296 2024-10-31 cs.LG cs.CL 57%

The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities

Venkatesh Balavadhani Parthasarathy, Ahtsham Zafar, Aafaq Khan, Arsalan Shahid

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19872 2024-10-29 cs.CV 57%

Radar and Camera Fusion for Object Detection and Tracking: A Comprehensive Survey

Kun Shi, Shibo He, Zhenyu Shi, Anjun Chen, Zehui Xiong, Jiming Chen, Jun Luo

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20684 2024-10-24 cs.SE cs.CL 57%

Joint Embeddings for Graph Instruction Tuning

Aaron Haag, Vlad Argatu, Oliver Lohse

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08350 2024-10-24 cs.CV 57%

CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model

Cheng Chen, Junchen Zhu, Xu Luo, Hengtao Shen, Lianli Gao, Jingkuan Song

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14049 2024-10-21 cs.CL 57%

Learning Metadata-Agnostic Representations for Text-to-SQL In-Context Example Selection

Chuhong Mai, Ro-ee Tal, Thahir Mohamed

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

Comments Accepted to NeurIPS 2024 Table Representation Learning workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13196 2024-10-21 cs.AI cs.LG 57%

Context-Enhanced Multi-View Trajectory Representation Learning: Bridging the Gap through Self-Supervised Models

Tangwen Qian, Junhe Li, Yile Chen, Gao Cong, Tao Sun, Fei Wang, Yongjun Xu

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13257 2024-10-18 cs.LG cs.AI 57%

scFusionTTT: Single-cell transcriptomics and proteomics fusion with Test-Time Training layers

Dian Meng, Bohao Xing, Xinlei Huang, Yanran Liu, Yijun Zhou, Yongjun xiao, Zitong Yu, Xubin Zheng

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏