$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
机构 * Tredence Inc.(特伦德公司) ; Indian Institute of Technology Kharagpur(印度理工学院克拉格浦尔分校) ; Inria Paris-Rocquencourt(巴黎-罗克琴库特研究所) ; Rajiv Gandhi University(拉吉夫·甘地大学) ; Tsinghua University(清华大学) ; Palmer Research Laboratories(帕勒姆研究实验室)
专题命中 评测与基准 :language model(title,abstract);分类 cs.CL
Comments 7 pages, 5 figures, 4 tables