Let's CONFER: A Dataset for Evaluating Natural Language Inference Models on CONditional InFERence and Presupposition
Tara Azin, Daniel Dumitrescu, Diana Inkpen, Raj Singh
机构
*
Carleton University(卡尔顿大学)
;
University of Ottawa(渥太华大学)
专题命中
推理评测
:reasoning(abstract);分类 cs.CL
CommentsThis paper is published in the Proceedings of the 38th Canadian Conference on Artificial Intelligence (CAIAC 2025). Please cite the conference version at https://caiac.pubpub.org/pub/keh8ij01
Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset
Diana Galvan-Sosa, Gabrielle Gaudeau, Pride Kavumba, Yunmeng Li, Hongyi gu, Zheng Yuan, Keisuke Sakaguchi, Paula Buttery
机构
*
ALTA Institute, Computer Laboratory, University of Cambridge(ALTA研究所、计算机实验室、剑桥大学)
;
SB Intuitions
;
Tohoku University(东北大学)
;
RIKEN(日本理化学研究所)
;
The University of Sheffield(谢菲尔德大学)
专题命中
推理评测
:reasoning(abstract);分类 cs.CL
Comments10 main pages (24 appendix pages), 9 figures, accepted to ACL 2025
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
The Hong Kong University of Science and Technology(香港科技大学)
;
University of California, Merced(加州大学梅德福分校)
;
ETH Zurich(苏黎世联邦理工学院)
;
City University of Hong Kong(香港城市大学)
;
Huawei Noah’s Ark Lab(华为诺亚实验室)
;
Sun Yat-sen University(中山大学)
;
MBZUAI
;
Chongqing University(重庆大学)
专题命中
推理评测
:reasoning(abstract);分类 cs.LG
Journal refThe Thirteenth International Conference on Learning Representations, 2025
Machine vs Machine: Using AI to Tackle Generative AI Threats in Assessment
Mohammad Saleh Torkestani, Taha Mansouri
专题命中
推理评测
:reasoning(abstract);分类 cs.AI
CommentsPaper presented at the Learning, Teaching & Student Experience 2025 Conference. The Chartered Association of Business Schools (CABS), Nottingham, UK
Marlon Tobaben, Mohamed Ali Souibgui, Rubèn Tito, Khanh Nguyen, Raouf Kerkouche, Kangsoo Jung, Joonas Jälkö, Lei Kang, Andrey Barsky, Vincent Poulain d'Andecy, Aurélie Joseph, Aashiq Muhamed, Kevin Kuo, Virginia Smith, Yusuke Yamasaki, Takumi Fukami, Kenta Niwa, Iifan Tyou, Hiro Ishii, Rio Yokota, Ragul N, Rintu Kutum, Josep Llados, Ernest Valveny, Antti Honkela, Mario Fritz, Dimosthenis Karatzas
机构
*
University of Helsinki(赫尔辛基大学)
;
Computer Vision Center, Universitat Autònoma de Barcelona(巴塞罗那自治大学计算机视觉中心)
;
CISPA Helmholtz Center for Information Security(信息安全赫尔姆霍茨中心)
;
INRIA(法国国家信息与自动化技术研究所)
;
Yooz
;
Carnegie Mellon University(卡内基梅隆大学)
;
NTT(日本NTT公司)
;
Institute of Science Tokyo(东京科学研究所)
;
Department of Computer Science(计算机科学系)
;
Mphasis AI & Applied Tech Lab at Ashoka, Ashoka University(阿什oka大学人工智能与应用技术实验室)
;
Koita Centre for Digital Health - Ashoka (KCDH-A)(阿什oka数字健康中心(KCDH-A))
;
Trivedi School of Biosciences, Ashoka University(阿什oka大学Trivedi生物科学学院)
MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book
Sau Lai Yip, Sunan He, Yuxiang Nie, Shu Pui Chan, Yilin Ye, Sum Ying Lam, Hao Chen
机构
*
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科学与技术大学)
;
Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(化学与生物工程系,香港科学与技术大学)
;
Division of Life Science, The Hong Kong University of Science and Technology(生命科学系,香港科学与技术大学)
机构
*
Pennsylvania State University(宾夕法尼亚州立大学)
;
Duke University(杜克大学)
;
University of Washington(华盛顿大学)
;
Oregon State University(俄勒冈州立大学)
;
Google DeepMind(谷歌DeepMind)
;
Nanyang Technological University(南洋理工大学)
;
Meta
;
AG2AI, Inc.(AG2AI公司)