Light Future: Multimodal Action Frame Prediction via InstructPix2Pix
机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen(数据科学学院,香港中文大学(深圳))
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
Comments 9 pages including appendix, 4 tables, 8 figures, to be submitted to WACV 2026