arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-02-02 至 2026-02-02 共收录 4 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 4 篇

2601.23253 2026-02-02 cs.CV cs.LG 81%

Training-Free Test-Time Adaptation with Brownian Distance Covariance in Vision-Language Models

无需训练的测试时适应与布朗距离协方差在视觉-语言模型中

Yi Zhang, Chun-Wun Cheng, Angelica I. Aviles-Rivero, Zhihai He, Liang-Jie Zhang

机构 * College of Computer Science and Software Engineering, Shenzhen University, China(深圳大学计算机科学与软件工程学院) University of Cambridge, Cambridge, UK(剑桥大学) Yau Mathematical Sciences Center, Tsinghua University, Beijing, China(清华大学尤里伊数学科学中心) Southern University of Science and Technology, Shenzhen, China(南方科技大学)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.LG

AI总结 TaTa通过布朗距离协方差实现无需训练的测试时适应,提升视觉-语言模型在领域偏移下的效率与稳定性,同时在泛化性能上取得突破。

Comments Accepted in ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19700 2026-02-02 cs.LG cs.AI 81%

Generalizable Multimodal Large Language Model Editing via Invariant Trajectory Learning

通过不变轨迹学习实现通用的多模态大语言模型编辑

Jiajie Su, Haoyuan Wang, Xiaohua Feng, Yunshan Ma, Xiaobo Xia, Yuyuan Li, Xiaolin Zheng, Jianmao Xiao, Chaochao Chen

机构 * Zhejiang University, China(浙江大学) Singapore Management University, Singapore(新加坡管理学院) National University of Singapore, Singapore(新加坡国立大学) Hangzhou Dianzi University, China(杭州电子科技大学) Jiangxi Normal University, China(江西师范大学)

专题命中 其他VLM :multimodal large language model(title);MLLM(abstract);分类 cs.AI、cs.LG

AI总结 本文提出ODEit框架,通过不变轨迹学习提升多模态大语言模型的编辑可靠性、局部性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21815 2026-02-02 cs.CY cs.AI cs.CL cs.SI 57%

Moral Outrage Shapes Commitments Beyond Attention: Multimodal Moral Emotions on YouTube in Korea and the US

道德愤怒影响承诺:韩国和美国YouTube上的多模态道德情感

Seongchan Park, Jaehong Kim, Hyeonseung Kim, Heejin Bin, Sue Moon, Wonjae Lee

机构 * KAIST(韩国科学技术院)

专题命中 其他VLM :vision language model(abstract);分类 cs.AI

AI总结 本研究通过多模态道德情感分类器分析YouTube上道德愤怒对用户参与度的影响,发现其在不同文化中均能提升观看和评论等互动行为。

Comments Accepted at The Web Conference 2026. We release Korean and English multimodal moral emotion classifiers

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22455 2026-02-02 cs.CV 57%

ScribbleSense: Generative Scribble-Based Texture Editing with Intent Prediction

ScribbleSense: 基于生成的涂鸦纹理编辑与意图预测

Yudi Zhang, Yeming Geng, Lei Zhang

机构 * School of Computer Science, Beijing Institute of Technology(计算机学院,北京理工大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

AI总结 ScribbleSense通过结合多模态大语言模型和图像生成模型,实现了基于生成的涂鸦纹理编辑与意图预测,提升交互式编辑性能。

Comments Accepted by IEEE TVCG. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏