arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-24 至 2025-09-24 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 3 篇

2509.18792 2025-09-24 cs.CL 57%

Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing

Sabri Boughorbel, Fahim Dalvi, Nadir Durrani, Majd Hawasly

机构 * Qatar Computing Research Institute, HBKU(卡塔尔计算研究所,哈瓦那大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments 12 pages, accepted to the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18132 2025-09-24 cs.AI 57%

Position Paper: Integrating Explainability and Uncertainty Estimation in Medical AI

Xiuyi Fan

机构 * Lee Kong Chian School of Medicine, College of Computing Data Science, Nanyang Technological University, Singapore

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments Accepted at the International Joint Conference on Neural Networks, IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18128 2025-09-24 cs.LG 57%

Accounting for Uncertainty in Machine Learning Surrogates: A Gauss-Hermite Quadrature Approach to Reliability Analysis

Amirreza Tootchi, Xiaoping Du

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏