arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-30 至 2025-09-30 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 10 篇

2509.23109 2025-09-30 cs.AI cs.CV 79%

AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors

Junyang Zhang, Tianyi Zhu, Thierry Tambe

机构 * California Institute of Technology(加州理工学院) Stanford University(斯坦福大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

Comments 31 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23002 2025-09-30 stat.ML cs.LG 79%

Unsupervised Conformal Inference: Bootstrapping and Alignment to Control LLM Uncertainty

Lingyou Pang, Lei Huang, Jianyu Lin, Tianyu Wang, Akira Horiguchi, Alexander Aue, Carey E. Priebe

机构 * Department of Statistics, University of California, Davis(加州大学戴维斯分校统计系) Department of Applied Mathematics and Statistics, Johns Hopkins University(约翰霍普金斯大学应用数学与统计学系)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.LG

Comments 26 pages including appendix; 3 figures and 5 tables. Under review for ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23665 2025-09-30 cs.LG cs.AI cs.IT math.IT math.PR 76%

Calibration Meets Reality: Making Machine Learning Predictions Trustworthy

Kristina P. Sinaga, Arjun S. Nair

机构 * Independent Researcher(独立研究者)

专题命中 幻觉与事实性 :trustworthy(title);分类 cs.AI、cs.LG

Comments 30 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09598 2025-09-30 cs.CL cs.AI cs.CY 67%

How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation

Ruohao Guo, Wei Xu, Alan Ritter

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23497 2025-09-30 cs.AI cs.HC cs.LG 62%

Dynamic Trust Calibration Using Contextual Bandits

Bruno M. Henrique, Eugene Santos

机构 * Thayer School of Engineering(泰勒工程学院) Dartmouth College(达特茅斯学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23146 2025-09-30 cs.CL cs.LG 62%

Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models

Zichao Yu, Ming Li, Wenyi Zhang, Weiguo Gao

机构 * University of Science and Technology of China(中国科学技术大学) Fudan University(复旦大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23088 2025-09-30 cs.CL 57%

The Geometry of Creative Variability: How Credal Sets Expose Calibration Gaps in Language Models

Esteban Garces Arias, Julian Rodemann, Christian Heumann

机构 * Department of Statistics, LMU Munich(统计系,慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) CISPA Helmholtz Center for Information Security(信息安全赫尔姆霍兹研究中心)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments Accepted at the 2nd UncertaiNLP Workshop @ EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07618 2025-09-30 cs.IR 50%

KAQG: A Knowledge-Graph-Enhanced RAG for Difficulty-Controlled Question Generation

Ching Han Chen, Ming Fang Shiu

专题命中 幻觉与事实性 :alignment(abstract)

Comments 10 pages, 4 figures and 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23457 2025-09-30 cs.CV 50%

No Concept Left Behind: Test-Time Optimization for Compositional Text-to-Image Generation

Mohammad Hossein Sameti, Amir M. Mansourian, Arash Marioriyad, Soheil Fadaee Oshyani, Mohammad Hossein Rohban, Mahdieh Soleymani Baghshah

机构 * Sharif University of Technology(沙里夫技术大学)

专题命中 幻觉与事实性 :alignment(abstract)

Comments 8 pages, 8 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23425 2025-09-30 cs.MA 50%

Situational Awareness for Safe and Robust Multi-Agent Interactions Under Uncertainty

Benjamin Alcorn, Eman Hammad

专题命中 幻觉与事实性 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏