arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-11 至 2025-12-11 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 5 篇

2512.09801 2025-12-11 cs.CV 83%

Modality-Specific Enhancement and Complementary Fusion for Semi-Supervised Multi-Modal Brain Tumor Segmentation

模态特异性增强与互补融合用于半监督多模态脑肿瘤分割

Tien-Dat Chung, Ba-Thinh Lam, Thanh-Huy Nguyen, Thien Nguyen, Nguyen Lan Vi Vu, Hoang-Loc Cao, Phat Kim Huynh, Min Xu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种半监督多模态脑肿瘤分割框架,通过模态特异性增强模块和互补信息融合模块提升分割性能。

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09311 2025-12-11 cs.CV cs.CR 79%

Transformer-Driven Multimodal Fusion for Explainable Suspiciousness Estimation in Visual Surveillance

基于Transformer的多模态融合用于视觉监控中的可解释性可疑性估计

Kuldeep Singh Yadav, Lalan Kumar

机构 * Big Data Research and Supercomputing Division, CSIR Fourth Paradigm Institute(CSIR第四范式研究所大数据研究与超级计算部门) Department of Electrical Engineering, Bharti School of Telecommunication, Yardi School of Artificial Intelligence, IIT Delhi(电信学院电子工程系、Yardi人工智能学院、德里理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于Transformer的多模态融合框架DeepUSEvision,结合YOLOv12、双深度卷积网络和Transformer判别器,实现高准确率和可解释性的可疑性估计。

Comments 12 pages, 10 figures, IEEE Transaction on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09010 2025-12-11 cs.CV cs.AI 73%

Towards Lossless Ultimate Vision Token Compression for VLMs

面向视觉语言模型的无损终极视觉标记压缩

Dehua Zheng, Mouxiao Huang, Borui Jiang, Hailin Hu, Xinghao Chen

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出LUVC框架,通过正交空间轴的迭代合并和频谱修剪单元,实现视觉语言模型中视觉标记的无损压缩,提升推理速度并保持精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06249 2025-12-11 cs.CL cs.AI 62%

TRepLiNa: Layer-wise CKA+REPINA Alignment Improves Low-Resource Machine Translation in Aya-23 8B

TRepLiNa:分层CKA+REPINA对齐改进Aya-23 8B低资源机器翻译

Toshiki Nakai, Ravi Kiran Chikkala, Lena Sophie Oberkircher, Nicholas Jennings, Natalia Skachkova, Tatiana Anikina, Jesujoba Oluwadara Alabi

机构 * Saarland University(萨尔兰大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 TRepLiNa通过结合CKA和REPINA实现分层对齐,提升低资源语言Aya-23 8B的机器翻译质量,尤其在数据稀缺情况下效果显著。

Comments It is work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09580 2025-12-11 cs.CV 57%

Content-Adaptive Image Retouching Guided by Attribute-Based Text Representation

基于属性文本表示的内容自适应图像润色

Hancheng Zhu, Xinyu Liu, Rui Yao, Kunyang Sun, Leida Li, Abdulmotaleb El Saddik

机构 * School of Computer Science and Technology/School of Artificial Intelligence, China University of Mining and Technology(计算机科学与技术学院/人工智能学院,中国矿业大学) Mine Digitization Engineering Research Center of the Ministry of Education, China University of Mining and Technology(教育部矿山数字化工程研究中心,中国矿业大学) School of Artificial Intelligence, Xidian University(人工智能学院,西安电子科技大学) School of Electrical Engineering and Computer Science, University of Ottawa(电气与计算机工程学院,多伦多大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出基于属性文本表示的内容自适应图像润色方法,通过结合内容感知的颜色调整和用户定义的风格偏好,实现高质量的图像润色效果。

详情

展开后加载摘要…

URL PDF HTML 收藏