Large Language Models are Strong Audio-Visual Speech Recognition Learners
专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted for publication at ICASSP 2025. The code and checkpoints are available here: https://github.com/umbertocappellazzo/Llama-AVSR