arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-30 至 2025-12-30 共收录 2 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 2 篇

2509.13754 2025-12-30 cs.CV 79%

Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval

跨模态全模式细粒度对齐用于文本到图像人物检索

Hao Yin, Xin Man, Feiyu Chen, Jie Shao, Heng Tao Shen

机构 * Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(深圳先进研究所,电子科学与技术大学) University of Electronic Science and Technology of China(电子科学与技术大学) Sichuan Artificial Intelligence Research Institute(四川人工智能研究院)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出FMFA框架,通过显式细粒度对齐和隐式关系推理实现文本到图像人物检索的高精度匹配。

Comments accepted by ACM Transactions on Multimedia Computing Communications and Applications in December 2025

Journal ref ACM Transactions on Multimedia Computing Communications and Applications, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22780 2025-12-30 cs.CV eess.IV 57%

Plug In, Grade Right: Psychology-Inspired AGIQA

插入选项,正确评分:心理学启发的AGIQA

Zhicheng Liao, Baoliang Chen, Hanwei Zhu, Lingyu Zhu, Shiqi Wang, Weisi Lin

机构 * School of Computer Science, South China Normal University(南方科技大学计算机科学学院) School of Computer Science and Engineering, Nanyang Technological University(南洋理工大学计算机科学与工程学院) School of Computer Science, City University of Hong Kong(香港城市大学计算机科学学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 基于心理学的AGIQA模型通过改进的分级响应模型提升图像质量评分的准确性与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏