Comments25 pages, 4 figures, 16 tables, 6 appendices. Code, task suite, released per-run verdicts, and a one-command reproduction of every reported number: https://github.com/shivenkk/agentrelbench
Comments13 pages, 4 figures. Accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE '26), Industry Showcase track, Munich, Germany, October 12-16, 2026
VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding
VideoGAIA:面向通用人工智能助手的智能体视频理解基准
Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu, Zheng Lian, Xinyu Geng, Jingdong Chen, Yi Yuan, Pheng-Ann Heng
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Ant Group(蚂蚁集团)
;
Tsinghua University(清华大学)
;
Tongji University(同济大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL
AnchorScore: A CLIP-Based Diagnostic of MLLM Annotation Difficulty
AnchorScore:一种基于CLIP的多模态大语言模型(MLLM)标注难度诊断方法
Yan Ma, Lizhuo Zhang
机构
*
School of Foreign Studies, Changsha University of Science and Technology(长沙理工大学外国语学院)
;
School of Education, Hunan Agricultural University(湖南农业大学教育学院)
;
School of Information and Intelligence, Hunan Agricultural University(湖南农业大学信息与智能科学学院)
专题命中
评测与基准
:large language model(abstract);language model(abstract)