arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-01 至 2025-08-01 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 3 篇

2507.22915 2025-08-01 cs.CL cs.AI 62%

Theoretical Foundations and Mitigation of Hallucination in Large Language Models

Esmail Gumaan

机构 * Department of Computer Science, University of Sana'a(桑加大学计算机科学系)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23736 2025-08-01 stat.ML cs.LG 57%

DICOM De-Identification via Hybrid AI and Rule-Based Framework for Scalable, Uncertainty-Aware Redaction

Kyle Naddeo, Nikolas Koutsoubis, Rahul Krish, Ghulam Rasool, Nidhal Bouaynaya, Tony OSullivan, Raj Krish

机构 * Rowan University(罗文大学) Moffitt Cancer Center(莫菲特癌症中心) University of South Florida(佛罗里达州立大学) Impact Business Information Solutions, Inc(Impact商务信息解决方案公司)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments 15 pages, 6 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20293 2025-08-01 cs.RO 50%

Decentralized Uncertainty-Aware Multi-Agent Collision Avoidance with Model Predictive Path Integral

Stepan Dergachev, Konstantin Yakovlev

机构 * FRC CSC RAS(俄罗斯科学院应用数学与控制论研究所) HSE University(俄罗斯高等经济大学) AIRI(人工智能研究所)

专题命中 幻觉与事实性 :safety(abstract)

Comments This is a pre-print of the paper accepted to IROS2025. The manuscript includes 8 pages, 4 figures, and 1 table. A supplementary video is available at https://youtu.be/_D4zDYJ4KCk Updated version: added link to source code in the abstract; updated experimental results description in Section VI.A; updated author affiliation and funding information; minor typo corrections

详情

展开后加载摘要…

URL PDF HTML 收藏