Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
Noy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope, Noam Slonim
机构
*
School of Computer Science and Engineering, The Hebrew University of Jerusalem(计算机科学与工程学院,特拉维夫大学)
;
IBM Research(IBM研究院)
;
The Allen Institute for AI (AI2)(人工智能研究所)
BeatFM: Improving Beat Tracking with Pre-trained Music Foundation Model
Ganghui Ru, Jieying Wang, Jiahao Zhao, Yulun Wu, Yi Yu, Nannan Jiang, Wei Wang, Wei Li
机构
*
School of Computer Science, Fudan University, Shanghai, China(复旦大学计算机科学学院)
;
Naval Medical Center, PLA, China(中国人民解放军海军医疗中心)
;
Graduate School of Informatics, Kyoto University, Kyoto, Japan(京都大学信息科学研究生院)
;
Graduate School of Advanced Science and Engineering, Hiroshima University, Hiroshima, Japan(广岛大学先进科学与工程研究生院)
;
Shanghai Key Laboratory of Intelligent Information Processing, Fudan University, Shanghai, China(上海智能信息处理重点实验室)
专题命中
评测与基准
:foundation model(title,abstract)
CommentsEarly draft for discussion only. Undergoing active revision, conclusions subject to change. Do not cite. Formal peer-reviewed version in preparation
Meaning-infused grammar: Gradient Acceptability Shapes the Geometric Representations of Constructions in LLMs
Supantho Rakshit, Adele Goldberg
机构
*
Princeton University(普林斯顿大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Comments6 pages, 3 figures, Accepted for publication at the Second International Workshop on Construction Grammars and NLP at the 16th International Conference for Computational Semantics (IWCS) 2025
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
Aisha Alansari, Hamzah Luqman
机构
*
Information and Computer Science Department, King Fahd University of Petroleum and Minerals(信息与计算机科学系,国王法赫德石油和矿物大学)
;
SDAIA-KFUPM Joint Research Center for Artificial Intelligence(SDAIA-KFUPM人工智能联合研究中心)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL
Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs
Manon Reusens, Bart Baesens, David Jurgens
机构
*
Research Centre for Information Systems Engineering (LIRIS), KU Leuven(信息系统工程研究中心(LIRIS),鲁汶大学)
;
Department of Engineering Management, University of Antwerp(工程管理系,安特卫普大学)
;
Department of Decision Analytics and Risk, University of Southampton(决策分析与风险系,南安普顿大学)
;
School of Information, University of Michigan(信息学院,密歇根大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL
FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain
Suifeng Zhao, Zhuoran Jin, Sujian Li, Jun Gao
机构
*
Key Laboratory of High Confidence Software Technologies, CS, Peking University, China(北京大学高可信软件技术重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Sciences, Peking University(北京大学多媒体信息处理国家重点实验室)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL