Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models
机构 * University of Virginia(弗吉尼亚大学) ; The Pennsylvania State University(宾夕法尼亚州立大学) ; Northeastern University(东北大学) ; Stanford University(斯坦福大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI