arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-01-28 至 2026-01-28 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 5 篇

2601.19267 2026-01-28 cs.CL 71%

DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models

DiaDem: 促进音频视频字幕中的对话描述以提升多模态大语言模型

Xinlong Chen, Weihong Lin, Jingyun Hua, Linli Yao, Yue Ding, Bozhou Li, Bohan Zeng, Yang Shi, Qiang Liu, Yuanxing Zhang, Pengfei Wan, Liang Wang, Tieniu Tan

机构 * New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Peking University(北京大学) Nanjing University(南京大学)

专题命中 其他VLM :multimodal large language model(title)

AI总结 DiaDem通过合成高质量数据集和难度分区的两阶段GRPO策略,提升了音频视频字幕中的对话描述准确性,并在多种基准测试中表现出色。

Comments Project webpage: https://diadem-captioner.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19659 2026-01-28 cs.CV cs.LG 62%

KeepLoRA: Continual Learning with Residual Gradient Adaptation

KeepLoRA: 基于残差梯度适应的持续学习

Mao-Lin Luo, Zi-Hao Zhou, Yi-Lin Zhang, Yuanyu Wan, Tong Wei, Min-Ling Zhang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China(教育部计算机网络与信息集成重点实验室) School of Software Technology, Zhejiang University(浙江大学软件学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.LG

AI总结 KeepLoRA通过残差梯度适应方法实现持续学习,有效平衡知识保留、任务知识保持和新知识获取,取得最佳性能。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22506 2026-01-28 cs.CR cs.AI 57%

SABRE-FL: Selective and Accurate Backdoor Rejection for Federated Prompt Learning

SABRE-FL:面向联邦提示学习的选样和准确后门拒绝

Momin Ahmad Khan, Yasra Chandio, Fatima Muhammad Anwar

机构 * University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

AI总结 SABRE-FL通过嵌入空间异常检测器有效识别并过滤联邦提示学习中的恶意客户端,显著降低后门攻击效果,提升系统鲁棒性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16435 2026-01-28 cs.CV cs.CL 57%

Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs

人类认知基准揭示大规模多模态语言模型中的基础视觉缺陷

Jen-Tse Huang, Dasen Dai, Jen-Yuan Huang, Youliang Yuan, Xiaoyuan Liu, Wenxuan Wang, Wenxiang Jiao, Pinjia He, Zhaopeng Tu, Haodong Duan

机构 * CUHK(香港中文大学) PKU(北京大学) CUHKSZ(香港中文大学深圳分校) RUC(中国人民大学) Tencent(腾讯) SHLab(深圳实验室)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

AI总结 人类认知基准揭示大规模多模态语言模型在基础视觉能力上的不足,通过VisFactor评估发现模型在关键视觉任务上表现不佳。

Comments Update: Evaluated 23 SOTA MLLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19203 2026-01-28 cs.HC 50%

Before Smelling the Video: A Two-Stage Pipeline for Interpretable Video-to-Scent Plans

在嗅觉之前:一种可解释的视频到香氛计划的两阶段流程

Kaicheng Wang, Kevin Zhongyang Shao, Ruiqi Chen, Sep Makhsous, Denise Wilson

专题命中 其他VLM :vision-language model(abstract)

AI总结 本文提出了一种两阶段流程,通过视觉语言模型和大型语言模型分离视频语义提取与气味推断,验证了语义规划在提升嗅觉媒体体验中的有效性。

Comments In submission of poster as ongoing project

详情

展开后加载摘要…

URL PDF HTML 收藏