Comments21 pages, 6 figures, 8 tables. Includes ancillary files with full benchmark results and ablation studies. Code available at https://github.com/athrael-soju/Snappy
机构
*
Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)
;
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
迈向基于自然语言的文档图像检索:新数据集和基准
Hao Guo, Xugong Qin, Jun Jie Ou Yang, Peng Zhang, Gangyan Zeng, Yubo Li, Hailun Lin
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Science and Engineering, Nanjing University of Science and Technology(南京理工大学 cyber 科学与工程学院)
;
State Key Laboratory of Cyberspace Security Defense(网络空间安全防御国家重点实验室)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
University of Southern California(美国南加州大学)
;
Laboratory for Advanced Computing and Intelligence Engineering(先进计算与智能工程实验室)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
LogicOCR: 大型多模态模型在文本丰富的图像上逻辑推理是否表现优异?
Maoyuan Ye, Haibin He, Qihuang Zhong, Jing Zhang, Juhua Liu, Bo Du
机构
*
School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University(计算机学院、多媒体软件国家工程研究中心、人工智能研究院、多媒体与网络通信工程湖北省重点实验室、武汉大学)
机构
*
Nankai University(南开大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
Wuhan University(武汉大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
Integrating Video and Text: A Balanced Approach to Multimodal Summary Generation and Evaluation
Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, Nancy F. Chen
机构
*
IRIT, University of Toulouse, France(IRIT,图卢兹大学,法国)
;
Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息通信研究所(I2R),A*STAR,新加坡)
;
CNRS, IRIT, France(CNRS,IRIT,法国)