AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
机构 * Tsinghua University(清华大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Fudan University(复旦大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 多模态评测 :omni-modal(title,abstract);multimodal(abstract);分类 cs.CV