TurkBench: A Benchmark for Evaluating Turkish Large Language Models
TurkBench:评估土耳其大型语言模型的基准测试
Çağrı Toraman, Ahmet Kaan Sever, Ayse Aysu Cengiz, Elif Ecem Arslan, Görkem Sevinç, Mete Mert Birdal, Yusuf Faruk Güldemir, Ali Buğra Kanburoğlu, Sezen Felekoğlu, Osman Gürlek, Sarp Kantar, Birsen Şahin Kütük, Büşra Tufan, Elif Genç, Serkan Coşkun, Gupse Ekin Demir, Muhammed Emin Arayıcı, Olgun Dursun, Onur Gungor, Susan Üsküdarlı, Abdullah Topraksoy, Esra Darıcı
When Noise Lowers The Loss: Rethinking Likelihood-Based Evaluation in Music Large Language Models
噪声降低损失:重新思考音乐大语言模型中的基于似然的评估
Xiaosha Li, Chun Liu, Ziyu Wang
机构
*
Georgia Institute of Technology(佐治亚理工学院)
;
ByteDance Inc.(字节跳动公司)
;
Courant Institute of Mathematical Sciences, New York University(纽约大学应用数学科学研究所)
;
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(Mohamed bin Zayed人工智能大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);分类 cs.AI
CommentsThis manuscript has been withdrawn by the authors because the methodology and results have been superseded by a more rigorous framework (SPACI and AST-ASIP). The corrected and expanded findings are now available in arXiv:2601.21360. Please cite the new manuscript instead
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models
评估大规模语言模型时代语言类型特征对多语言机器翻译的影响
Vitalii Hirak, Jaap Jumelet, Arianna Bisazza
机构
*
Data & Knowledge Engineering, Heinrich Heine University(海因里希·海因大学数据与知识工程系)
;
Center for Language and Cognition (CLCG), University of Groningen(格罗宁根大学语言与认知中心)
专题命中
评测与基准
:language model(title,abstract);large language model(title);分类 cs.CL
机构
*
Computational Modeling and Simulation University of Pittsburgh(计算建模与仿真大学匹兹堡大学)
;
Mathematics & Statistics Department University of Minnesota Duluth(数学与统计学系明尼苏达大学 Duluth分校)
;
Independent Researcher in AI and Statistics(人工智能与统计学独立研究者)
;
Hardware Technology Organization(硬件技术组织)
专题命中
评测与基准
:LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
AudioJailbreak: 对端到端大音频-语言模型的攻击
Guangke Chen, Fu Song, Zhe Zhao, Xiaojun Jia, Yang Liu, Yanchen Qiao, Weizhe Zhang, Weiping Tu, Yuhong Yang, Bo Du
机构
*
Wuhan University(武汉大学)
;
Key Laboratory of System Software (Chinese Academy of Sciences)(中国科学院系统软件重点实验室)
;
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
;
State Key Laboratory of Cryptology(密码学国家重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Nanjing Institute of Software Technology(南京软件技术研究所)
;
Ant Group(蚂蚁集团)
;
Nanyang Technological University(南洋理工大学)
;
Zhejiang Lab(浙江实验室)
;
Pengcheng Laboratory(鹏城实验室)
机构
*
1 Applied AI Institute, Moscow 121205 , Russia
;
(R.S.) 2 School of Computer Science, University of Lincoln, Lincoln LN6 7TS , UK 3 Independent Researcher, Dubai 500001, United Arab Emirates
TIDE: Trajectory-based Diagnostic Evaluation of Test-Time Improvement in LLM Agents
基于轨迹的测试时改进诊断评估:LLM代理中的测试时改进
Hang Yan, Xinyu Che, Fangzhi Xu, Qiushi Sun, Zichen Ding, Kanzhi Cheng, Jian Zhang, Tao Qin, Jun Liu, Qika Lin
机构
*
Xi’an Jiaotong University(西安交通大学)
;
The University of Hong Kong(香港大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Nanjing University(南京大学)
;
National University of Singapore(新加坡国立大学)
机构
*
School of Computer Science Chongqing University(重庆大学计算机学院)
;
MAIS Institute of Automation Chinese Academy of Sciences(中国科学院自动化研究所MAIS研究所)
;
The First Affiliated Hospital of Chongqing Medical University(重庆医科大学第一附属医院)
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI