arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-01-15 至 2026-01-15 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 3 篇

2601.08871 2026-01-15 cs.SD cs.AI eess.AS 79%

Semantic visually-guided acoustic highlighting with large vision-language models

语义视觉引导的音频突出与大型视觉-语言模型

Junhua Huang, Chao Huang, Chenliang Xu

机构 * University of Rochester(罗切斯特大学)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.AI

AI总结 本文提出利用大型视觉-语言模型提取视觉-语义特征,以提升音频混音质量,发现摄像机焦点、语气和场景背景对感知混音质量提升最显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15854 2026-01-15 cs.CV cs.LG 62%

Privacy-Preserving in Connected and Autonomous Vehicles Through Vision to Text Transformation

通过视觉到文本转换实现连接和自动驾驶车辆中的隐私保护

Abdolazim Rezaei, Mehdi Sookhak, Ahmad Patooghy, Shahab S. Band, Amir Mosavi

机构 * organization= Department of Computre Science, Texas A\&M University Corpus Christi , addressline= 6300 Ocean Dr , city= Corpus Christi , postcode= 78412 , state= Texas , country= USA organization= Department of Information Management, International Graduate School of Artificial Intelligence, National Yunlin University of Science organization= Institute of the Information Society, Ludovika University of Public Service , city= Budapest , postcode= 78412 , country= Hungary organization= North Carolina A\&T State University , addressline= 601 E Market , city= Greensboro , postcode= 27411 , country= USA

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.LG

AI总结 本文提出利用强化学习和视觉-语言模型,通过视觉到文本转换实现对连接和自动驾驶车辆中敏感视觉信息的隐私保护,提升了隐私保护效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09470 2026-01-15 physics.ed-ph cs.AI 57%

Personalized Multimodal Feedback Using Multiple External Representations: Strategy Profiles and Learning in High School Physics

基于多种外部表征的个性化反馈:策略配置与高中物理学习中的学习

Natalia Revenga-Lozano, Karina E. Avila, Steffen Steinert, Matthias Schweinberger, Clara E. Gómez-Pérez, Jochen Kuhn, Stefan Küchemann

机构 * Chair of Physics Education, Faculty of Physics, Ludwig-Maximilians-Universität München (LMU Munich)(物理教育系主任,物理学院,慕尼黑路易斯-马克西姆利安大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

AI总结 本文研究了多种外部表征与个性化反馈在高中物理学习中的整合效果,发现详细多表征反馈对学习成绩有积极影响,且学习者根据表征能力选择不同反馈策略。

Comments Keywords: Adaptive Feedback, Multimodal Learning, Multiple External Representations, Physics Education, Science Education, Representational Competences, Intelligent Tutoring Systems

详情

展开后加载摘要…

URL PDF HTML 收藏