arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-23 至 2025-10-23 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 3 篇

2510.19476 2025-10-23 cs.LG cs.AI 81%

A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring

Julian Schulz

机构 * Meridian Research, Cambridge(梅迪安研究,剑桥)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18918 2025-10-23 cs.CL cs.AI 62%

Misinformation Detection using Large Language Models with Explainability

Jainee Patel, Chintan Bhatt, Himani Trivedi, Thanh Thi Nguyen

机构 * Department of Computer Engineering, LDRP Institute of Technology and Research(计算机工程系,LDRP技术与研究学院) University of Wollongong(沃林根大学) Monash University(莫纳什大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication in the Proceedings of the 8th International Conference on Algorithms, Computing and Artificial Intelligence (ACAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19310 2025-10-23 cs.CL 57%

JointCQ: Improving Factual Hallucination Detection with Joint Claim and Query Generation

Fan Xu, Huixuan Zhang, Zhenliang Zhang, Jiahao Wang, Xiaojun Wan

机构 * Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学) Trustworthy Technology and Engineering Laboratory, Huawei(可信技术与工程实验室,华为)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏