arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

2026-07-20 至 2026-07-20 共收录 3
2607.16094 2026-07-20 cs.CV 新提交

How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA

视觉语言模型如何失败?组合式视觉问答中的视觉-操作不对齐

Navya Gupta, Bingjie Xu, Avinash Anand, Timothy Liu, Zhengchen Zhang

机构 * Singapore Institute of Technology(新加坡科技学院) NVIDIA(英伟达)

AI总结 研究组合式视觉问答中视觉语言模型失败的机制,引入以操作为中心的框架分解失败模式,揭示四种失败模式及传播路径,表明不同失败类型需不同纠正策略,为提升模型可靠性提供基础。

Comments Accepted at ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15768 2026-07-20 cs.CV cs.AI 新提交

GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

GeoChrono:遥感中长期时间理解的基准测试与重新思考

Yujie Li, Jiancheng Pan, Zhiwei Wei, Jiuniu Wang, Mugen Peng, Wenjia Xu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Hunan Normal University(湖南师范大学) City University of Hong Kong(香港城市大学)

AI总结 研究针对遥感中长期时间理解缺乏系统评估的问题,引入ChronoBench基准。基于评估结果提出GeoChrono模型,设计相关编码器和压缩器并构建训练数据集,该模型在基准测试中性能领先,压缩器有效减少视觉令牌。

Comments Accepted to ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15752 2026-07-20 cs.CV 新提交

Personalized Image Aesthetic Assessment via Preference-rich Sample Mining and Cohort Merging

通过偏好丰富样本挖掘和群组合并进行个性化图像美学评估

Zhichao Yang, Tianjiao Gu, Zhixianhe Zhang, Xiangfei Sheng, Pengfei Chen, Leida Li

机构 * Xidian University(西安电子科技大学)

AI总结 研究个性化图像美学评估问题,提出基于多模态大语言模型的PRAC方法,通过偏好丰富样本挖掘和美学共鸣群组合并建模个体美学偏好,经实验验证该方法优于现有技术。

Comments The paper has been accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏