arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2412.02294 2024-12-04 eess.IV cs.AI cs.CV cs.LG 62%

Initial Study On Improving Segmentation By Combining Preoperative CT And Intraoperative CBCT Using Synthetic Data

Maximilian E. Tschuchnig, Philipp Steininger, Michael Gadermayr

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at BVM 2025. arXiv admin note: text overlap with arXiv:2406.11650

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14811 2024-12-03 cs.CV cs.CL cs.LG 62%

Fine-Grained Alignment in Vision-and-Language Navigation through Bayesian Optimization

Yuhang Song, Mario Gianni, Chenguang Yang, Kunyang Lin, Te-Chuan Chiu, Anh Nguyen, Chun-Yi Lee

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16824 2024-11-27 cs.CV cs.AI cs.LG 62%

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge

Yaqi Zhao, Yuanyang Yin, Lin Li, Mingan Lin, Victor Shea-Jay Huang, Siwei Chen, Weipeng Chen, Baoqun Yin, Zenan Zhou, Wentao Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02684 2024-11-26 cs.CV cs.CL 62%

Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE

Zeren Chen, Ziqin Wang, Zhen Wang, Huayang Liu, Zhenfei Yin, Si Liu, Lu Sheng, Wanli Ouyang, Yu Qiao, Jing Shao

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 22 pages, 12 figures. Accepted in ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19581 2024-11-22 cs.SE cs.AI cs.CL 62%

Source Code Foundation Models are Transferable Binary Analysis Knowledge Bases

Zian Su, Xiangzhe Xu, Ziyang Huang, Kaiyuan Zhang, Xiangyu Zhang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07516 2024-11-13 cs.CV cs.CL 62%

SparrowVQE: Visual Question Explanation for Course Content Understanding

Jialu Li, Manish Kumar Thota, Ruslan Gokhman, Radek Holik, Youshan Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04308 2024-11-08 cs.CL cs.AI 62%

Improving Bilingual Capabilities of Language Models to Support Diverse Linguistic Practices in Education

Anand Syamkumar, Nora Tseng, Kaycie Barron, Shanglin Yang, Shamya Karumbaiah, Rheeya Uppal, Junjie Hu

专题命中 多模态训练与对齐 :MLLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01723 2024-11-08 cs.CV cs.AI 62%

Towards Calibrated Robust Fine-Tuning of Vision-Language Models

Changdae Oh, Hyesu Lim, Mijoo Kim, Dongyoon Han, Sangdoo Yun, Jaegul Choo, Alexander Hauptmann, Zhi-Qi Cheng, Kyungwoo Song

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2024 (a short version was presented at the NeurIPS 2023 Workshop on Distribution Shifts); Major modification of (v7): Fixing the x-axis of Figure 3 and Pearson correlation, accordingly

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00860 2024-11-05 cs.CL cs.CV 62%

Survey of Cultural Awareness in Language Models: Text and Beyond

Siddhesh Pawar, Junyeong Park, Jiho Jin, Arnav Arora, Junho Myung, Srishti Yadav, Faiz Ghifari Haznitrama, Inhwa Song, Alice Oh, Isabelle Augenstein

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00477 2024-11-04 cs.SD cs.AI eess.AS q-bio.QM 62%

Multi Modal Information Fusion of Acoustic and Linguistic Data for Decoding Dairy Cow Vocalizations in Animal Welfare Assessment

Bubacarr Jobarteh, Madalina Mincu, Gavojdian Dinu, Suresh Neethirajan

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI、eess.AS

Comments 31 pages, 22 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17510 2024-10-29 q-bio.NC cs.AI cs.CV cs.LG 62%

NeuroPath: A Neural Pathway Transformer for Joining the Dots of Human Connectomes

Ziquan Wei, Tingting Dan, Jiaqi Ding, Guorong Wu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17270 2024-10-24 cs.AI cs.CL cs.LG cs.LO cs.NE 62%

Proof of Thought : Neurosymbolic Program Synthesis allows Robust and Interpretable Reasoning

Debargha Ganguly, Srinivasan Iyengar, Vipin Chaudhary, Shivkumar Kalyanaraman

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 38th Conference on Neural Information Processing Systems (NeurIPS 2024) System 2 Reasoning At Scale Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15926 2024-10-22 cs.CV cs.CL 62%

Mitigating Object Hallucination via Concentric Causal Attention

Yun Xing, Yiheng Li, Ivan Laptev, Shijian Lu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

Comments To appear at NeurIPS 2024. Code is available at https://github.com/xing0047/cca-llava

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13351 2024-10-18 cs.CL cs.AI cs.LG 62%

Representation Learning of Structured Data for Medical Foundation Models

Vijay Prakash Dwivedi, Viktor Schlegel, Andy T. Liu, Thanh-Tung Nguyen, Abhinav Ramesh Kashyap, Jeng Wei, Wei-Hsian Yin, Stefan Winkler, Robby T. Tan

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2024 Workshop on Unifying Representations in Neural Models (UniReps 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11657 2024-10-16 cs.CL cs.CV 62%

Unveiling the Mystery of Visual Attributes of Concrete and Abstract Concepts: Variability, Nearest Neighbors, and Challenging Categories

Tarun Tater, Sabine Schulte im Walde, Diego Frassinelli

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17942 2024-10-15 cs.CV cs.AI cs.RO 62%

Learning Shared RGB-D Fields: Unified Self-supervised Pre-training for Label-efficient LiDAR-Camera 3D Perception

Xiaohao Xu, Ye Li, Tianyi Zhang, Jinrong Yang, Matthew Johnson-Roberson, Xiaonan Huang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09047 2024-10-14 cs.CL cs.AI cs.LG 62%

Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models

Qin Liu, Chao Shang, Ling Liu, Nikolaos Pappas, Jie Ma, Neha Anna John, Srikanth Doss, Lluis Marquez, Miguel Ballesteros, Yassine Benajiba

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05997 2024-10-10 eess.AS cs.CV cs.LG cs.SD 62%

An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment

Hugo Malard, Michel Olvera, Stéphane Lathuiliere, Slim Essid

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05650 2024-10-10 cs.CV cs.MM 62%

SIA-OVD: Shape-Invariant Adapter for Bridging the Image-Region Gap in Open-Vocabulary Detection

Zishuo Wang, Wenhao Zhou, Jinglin Xu, Yuxin Peng

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV、cs.MM

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04038 2024-10-10 cs.AI cs.CV 62%

Gamified crowd-sourcing of high-quality data for visual fine-tuning

Shashank Yadav, Rohan Tomar, Garvit Jain, Chirag Ahooja, Shubham Chaudhary, Charles Elkan

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04751 2024-10-08 cs.CV cs.CL 62%

Intriguing Properties of Large Language and Vision Models

Young-Jun Lee, Byungsoo Ko, Han-Gyu Kim, Yechan Hwang, Ho-Jin Choi

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Code is available in https://github.com/passing2961/IP-LLVM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02271 2024-10-04 cs.SD cs.AI eess.AS 62%

CoLLAP: Contrastive Long-form Language-Audio Pretraining with Musical Temporal Structure Augmentation

Junda Wu, Warren Li, Zachary Novack, Amit Namburi, Carol Chen, Julian McAuley

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19603 2024-10-01 cs.CV cs.AI 62%

One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Zechen Bai, Tong He, Haiyang Mei, Pichao Wang, Ziteng Gao, Joya Chen, Lei Liu, Zheng Zhang, Mike Zheng Shou

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurlPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08044 2024-09-30 cs.CL cs.AI cs.LG 62%

RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization

Xijie Huang, Zechun Liu, Shih-Yang Liu, Kwang-Ting Cheng

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024 Findings, Codes: https://github.com/HuangOwen/RoLoRA, Models: https://huggingface.co/collections/ScarletAce/rolora-66f5f228a90681c7c4512b28

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18576 2024-09-27 cs.CV cs.AI 62%

Fixed-length Dense Descriptor for Efficient Fingerprint Matching

Zhiyu Pan, Yongjie Duan, Jianjiang Feng, Jie Zhou

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by WIFS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.20058 2024-09-23 eess.IV cs.AI cs.CV cs.LG 62%

Revolutionizing Disease Diagnosis with simultaneous functional PET/MR and Deeply Integrated Brain Metabolic, Hemodynamic, and Perfusion Networks

Luoyu Wang, Yitian Tao, Qing Yang, Yan Liang, Siwei Liu, Hongcheng Shi, Dinggang Shen, Han Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12716 2024-09-20 cs.CV cs.AI 62%

Optical Flow Matters: an Empirical Comparative Study on Fusing Monocular Extracted Modalities for Better Steering

Fouad Makiyeh, Mark Bastourous, Anass Bairouk, Wei Xiao, Mirjana Maras, Tsun-Hsuan Wangb, Marc Blanchon, Ramin Hasani, Patrick Chareyre, Daniela Rus

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11375 2024-09-18 cs.CV cs.AI 62%

Multi-OCT-SelfNet: Integrating Self-Supervised Learning with Multi-Source Data Fusion for Enhanced Multi-Class Retinal Disease Classification

Fatema-E- Jannat, Sina Gholami, Jennifer I. Lim, Theodore Leng, Minhaj Nur Alam, Hamed Tabkhi

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 25 pages, 9 tables, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10164 2024-09-17 cs.LG cs.AI cs.CL 62%

Quantile Regression for Distributional Reward Models in RLHF

Nicolai Dorka

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14950 2024-08-28 cs.CV cs.AI 62%

NeuralOOD: Improving Out-of-Distribution Generalization Performance with Brain-machine Fusion Learning Framework

Shuangchen Zhao, Changde Du, Hui Li, Huiguang He

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏