TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health
TrustMH-Bench: 一个用于评估大语言模型在心理健康领域可信度的综合基准
Zixin Xiong, Ziteng Wang, Haotian Fan, Xinjie Zhang, Wenxuan Wang
机构
*
Renmin University of China, China(中国人民大学)
;
Beijing University of Posts and Telecommunications, China(北京邮电大学)
;
Hefei University of Technology, China(合肥工业大学)
PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology
PanCanBench: 用于评估大型语言模型在胰腺肿瘤学中的综合基准
Yimin Zhao, Sheela R. Damle, Simone E. Dekker, Scott Geng, Karly Williams Silva, Jesse J Hubbard, Manuel F Fernandez, Fatima Zelada-Arenas, Alejandra Alvarez, Brianne Flores, Alexis Rodriguez, Stephen Salerno, Carrie Wright, Zihao Wang, Pang Wei Koh, Jeffrey T. Leek
机构
*
Department of Biostatistics, University of Washington(华盛顿大学生物统计学系)
;
Clinical Research Division, Fred Hutch Cancer Center(Fred Hutch癌症中心临床研究部)
;
Division of Hematology and Oncology, Department of Medicine, University of Washington(华盛顿大学医学系血液学与肿瘤学分会)
;
Allen Institute for AI(Allen人工智能研究所)
;
Department of Computer Science and Engineering, University of Washington(华盛顿大学计算机科学与工程系)
;
Public Health Sciences, Biostatistics, Fred Hutchinson Cancer Center(Fred Hutchinson癌症中心公共卫生科学与生物统计学)
Capabilities Ain't All You Need: Measuring Propensities in AI
能力并非全部所需:测量AI倾向性
Daniel Romero-Alvarado, Fernando Martínez-Plumed, Lorenzo Pacchiardi, Hugo Save, Siddhesh Milind Pawar, Behzad Mehrbakhsh, Pablo Antonio Moreno Casares, Ben Slater, Paolo Bova, Peter Romero, Zachary R. Tidler, Jonathan Prunty, Luning Sun, Jose Hernandez-Orallo
机构
*
Valencian Research Institute of Artificial Intelligence, Universitat Politècnica de València, Valencia, Spain
;
University of Copenhagen, Denmark work done while at University of Cambridge
;
Existential Risk Observatory, Amsterdam, Netherlands
;
Leverhulme Centre for the Future of Intelligence, University of Cambridge
;
The Psychometrics Centre, University of Cambridge
;
Department of Computing \& Games, University of Teesside
;
Georgia Institute of Technology
;
University of Cambridge
机构
*
Zhejiang University(浙江大学)
;
Binjiang Institute of Zhejiang University(浙江大学滨江学院)
;
Om AI Research(奥姆人工智能研究)
;
Stanford University(斯坦福大学)
;
ETH Zürich(苏黎世联邦理工学院)
Crash Severity Risk Modeling Strategies under Data Imbalance
在数据不平衡情况下工作区碰撞严重性风险建模策略
Abdullah Al Mamun, Abyad Enan, Debbie A. Indah, Judith Mwakalonge, Gurcan Comert, Mashrur Chowdhury
专题命中
安全评测
:safety(abstract);分类 cs.CY、cs.LG
AI总结
本研究探讨了在数据不平衡情况下,利用DMI特征选择和不同模型提升HS碰撞预测性能的方法。
CommentsThis second revised version has been resubmitted to the Transportation Research Record: Journal of the Transportation Research Board after addressing the reviewers' comments and is currently awaiting the final decision
Interpolation-Driven Machine Learning Approaches for Plume Shine Dose Estimation: A Comparison of XGBoost, Random Forest, and TabNet
基于插值驱动的机器学习方法用于烟云闪烁剂量估计:XGBoost、随机森林和TabNet的比较
Biswajit Sadhu, Kalpak Gupte, Trijit Sadhu, S. Anand
机构
*
Health Physics Division, Health Safety \& Environment Group, Bhabha Atomic Research Center, Mumbai – 400085, India
;
Birla Institute of Technology And Science, Pilani, Rajasthan – 333031, India
;
Environment Group, Bhabha Atomic Research Center, Mumbai – 400085, India
CommentsThis work was submitted without the consent of my current adviser. Additionally, it overlaps with my unpublished research work. In order to avoid potential academic and authorship conflicts, I am requesting withdrawal of the paper