arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-07 至 2025-10-07 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 7 篇

2510.04045 2025-10-07 cs.CL cs.LG 81%

Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment

Yunfan Zhang, Kathleen McKeown, Smaranda Muresan

机构 * Columbia University(哥伦比亚大学) Barnard College(巴纳德学院)

专题命中 幻觉与事实性 :alignment(title);safety(abstract);分类 cs.CL、cs.LG

Comments ACL EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17225 2025-10-07 cs.CL cs.AI 73%

SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation

Xiaqiang Tang, Yi Wang, Keyu Hu, Rui Xu, Chuang Li, Weigao Sun, Jian Li, Sihong Xie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Chinese Academy of Sciences(中国科学院) Shanghai AI Lab(上海人工智能实验室) Tencent Hunyuan(腾讯文英)

专题命中 幻觉与事实性 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04933 2025-10-07 cs.CL cs.AI cs.IT cs.LG cs.NE math.IT 67%

The Geometry of Truth: Layer-wise Semantic Dynamics for Hallucination Detection in Large Language Models

Amir Hameed Mir

机构 * Sirraya Labs(Sirraya实验室)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Comments: 14 pages, 14 figures, 5 tables. Code available at: https://github.com/sirraya-tech/Sirraya_LSD_Code

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04439 2025-10-07 cs.CL 57%

On the Role of Unobserved Sequences on Sample-based Uncertainty Quantification for LLMs

Lucie Kunitomo-Jacquin, Edison Marrese-Taylor, Ken Fukuda

机构 * National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments Accepted to UncertaiNLP workshop of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04357 2025-10-07 cs.LG q-fin.CP 57%

From News to Returns: A Granger-Causal Hypergraph Transformer on the Sphere

Anoushka Harit, Zhongtian Sun, Jongmin Yu

机构 * University of Cambridge(剑桥大学) University of Kent(肯特大学) Department of Computer Science, University of Cambridge(剑桥大学计算机科学系)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments 6th ACM International Conference on AI in Finance

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10246 2025-10-07 cs.LG 57%

Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions

Hazel Kim, Tom A. Lamb, Adel Bibi, Philip Torr, Yarin Gal

机构 * University of Oxford(牛津大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

Comments Accepted to EMNLP(main)2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20148 2025-10-07 cs.CV 50%

Smaller is Better: Enhancing Transparency in Vehicle AI Systems via Pruning

Sanish Suwal, Shaurya Garg, Dipkamal Bhusal, Michael Clifford, Nidhi Rastogi

专题命中 幻觉与事实性 :safety(abstract)

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏