arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-20 至 2025-11-20 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 3 篇

2503.15511 2025-11-20 cs.HC cs.CY cs.LG 62%

The Trust Calibration Maturity Model for Characterizing and Communicating Trustworthiness of AI Systems

Scott T Steinmetz, Asmeret Naugle, Paul Schutte, Matt Sweitzer, Alex Washburne, Lisa Linville, Daniel Krofcheck, Michal Kucer, Samuel Myren

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CY、cs.LG

Comments 19 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15005 2025-11-20 cs.CL cs.AI 62%

Mathematical Analysis of Hallucination Dynamics in Large Language Models: Uncertainty Quantification, Advanced Decoding, and Principled Mitigation

Moses Kiprono

机构 * Catholic University of America(美国天主教大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments 10 pages, theoretical/mathematical LLM research, no figures, intended for peer-reviewed journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01840 2025-11-20 cs.LG 57%

Optimizing In-Context Learning for Efficient Full Conformal Prediction

Weicao Deng, Sangwoo Park, Min Li, Osvaldo Simeone

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏