arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1755 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1755 篇

2511.01902 2025-11-05 cs.CY cs.AI 62%

Before the Clinic: Transparent and Operable Design Principles for Healthcare AI

Alexander Bakumenko, Aaron J. Masino, Janine Hoelscher

机构 * Clemson University(克莱姆森大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15189 2025-10-30 cs.CL cs.CY 62%

Face the Facts! Evaluating RAG-based Pipelines for Professional Fact-Checking

Daniel Russo, Stefano Menini, Jacopo Staiano, Marco Guerini

机构 * Fondazione Bruno Kessler(布罗诺·凯塞勒基金会) University of Trento(特伦托大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.CY

Comments Code and data at https://github.com/drusso98/face-the-facts - Accepted for publication at INLG 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23264 2025-10-28 cs.LG cs.AI 62%

PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization

Xinhai Wang, Shu Yang, Liangyu Wang, Lin Zhang, Huanyi Xie, Lijie Hu, Di Wang

机构 * King Abdullah University of Science and Technology(卡布斯大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22261 2025-10-28 cs.LG cs.AI 62%

Epistemic Deep Learning: Enabling Machine Learning Models to Know When They Do Not Know

Shireen Kudukkil Manchingal

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22751 2025-10-28 cs.AI cs.CL 62%

Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models

Piyushkumar Patel

机构 * Microsoft(微软)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22362 2025-10-28 cs.LG cs.CL 62%

Mapping Faithful Reasoning in Language Models

Jiazheng Li, Andreas Damianou, J Rosser, José Luis Redondo García, Konstantina Palla

机构 * King’s College London(伦敦国王学院) Spotify UK(Spotify英国分公司) University of Oxford(牛津大学) Spotify Spain(Spotify西班牙分公司)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.LG

Comments 9 pages, Accepted to the Mechanistic Interpretability Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06771 2025-10-27 cs.AI cs.CV cs.LG 62%

Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty

Meera Hahn, Wenjun Zeng, Nithish Kannen, Rich Galt, Kartikeya Badola, Been Kim, Zi Wang

机构 * Google DeepMind(谷歌DeepMind)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

Journal ref International Conference on Machine Learning, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18918 2025-10-23 cs.CL cs.AI 62%

Misinformation Detection using Large Language Models with Explainability

Jainee Patel, Chintan Bhatt, Himani Trivedi, Thanh Thi Nguyen

机构 * Department of Computer Engineering, LDRP Institute of Technology and Research(计算机工程系,LDRP技术与研究学院) University of Wollongong(沃林根大学) Monash University(莫纳什大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication in the Proceedings of the 8th International Conference on Algorithms, Computing and Artificial Intelligence (ACAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15233 2025-10-20 cs.LG cs.AI 62%

Adaptive Individual Uncertainty under Out-Of-Distribution Shift with Expert-Routed Conformal Prediction

Amitesh Badkul, Lei Xie

机构 * Ph.D. Programs in Computer Science(计算机科学博士项目) The Graduate Center, City University of New York(纽约城市大学研究生中心) School of Pharmacy and Pharmaceutical Sciences(药学与制药科学学院) Center for Drug Discovery(药物发现中心) Northeastern University(东北大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01997 2025-10-17 cs.LG cs.AI stat.ML 62%

Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach

Jiancong Xiao, Bojian Hou, Zhanliang Wang, Ruochen Jin, Qi Long, Weijie J. Su, Li Shen

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

Journal ref ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19234 2025-10-16 cs.AI cs.CL cs.MA 62%

GUARDIAN: Safeguarding LLM Multi-Agent Collaborations with Temporal Graph Modeling

Jialong Zhou, Lichao Wang, Xiao Yang

机构 * King’s College London(伦敦国王学院) Beijing Institute of Technology(北京理工大学) Tsinghua University(清华大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12324 2025-10-08 cs.CL cs.AI 62%

Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Prediction

Mengying Yuan, Wenhao Wang, Zixuan Wang, Yujie Huang, Kangli Wei, Fei Li, Chong Teng, Donghong Ji

机构 * Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航空信息安全与可信计算重点实验室,教育部,网络安全与工程学院,武汉大学) Zhejiang University(浙江大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 Main (Camera Ready)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02571 2025-10-06 cs.CV cs.AI cs.CL 62%

How Confident are Video Models? Empowering Video Models to Express their Uncertainty

Zhiting Mei, Ola Shorinwa, Anirudha Majumdar

机构 * Princeton University(普林斯顿大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01288 2025-10-03 cs.LG cs.AI 62%

Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours

Rui Melo, Rui Abreu, Corina S. Pasareanu

机构 * Carnegie Mellon University(卡内基梅隆大学) FEUP(费拉尔大学) INESC-ID(葡萄牙里斯本信息技术与创新研究中心)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments 9 main pages, 13 appendix pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01237 2025-10-03 cs.CL cs.AI 62%

Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation

Nandakishor M

机构 * AI Safety Research(人工智能安全研究) Convai Innovations(Convai创新)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23585 2025-10-02 cs.LG cs.AI cs.CV 62%

EVO-LRP: Evolutionary Optimization of LRP for Interpretable Model Explanations

Emerald Zhang, Julian Weaver, Samantha R Santacruz, Edward Castillo

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23497 2025-09-30 cs.AI cs.HC cs.LG 62%

Dynamic Trust Calibration Using Contextual Bandits

Bruno M. Henrique, Eugene Santos

机构 * Thayer School of Engineering(泰勒工程学院) Dartmouth College(达特茅斯学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23146 2025-09-30 cs.CL cs.LG 62%

Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models

Zichao Yu, Ming Li, Wenyi Zhang, Weiguo Gao

机构 * University of Science and Technology of China(中国科学技术大学) Fudan University(复旦大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19375 2025-09-25 cs.LG cs.AI stat.ML 62%

Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation

Mridul Sharma, Adeetya Patel, Zaneta D' Souza, Samira Abbasgholizadeh Rahimi, Siva Reddy, Sreenath Madathil

机构 * Faculty of Dental Medicine and Oral Health Sciences, McGill University(牙医学院与口腔健康科学学院,麦吉尔大学) McGill University(麦吉尔大学) Mila–Quebec Artificial Intelligence Institute(魁北克人工智能研究所) School of Computer Science, McGill University(计算机科学学院,麦吉尔大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17671 2025-09-23 cs.CL cs.AI 62%

Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications

Selva Taş, Mahmut El Huseyni, Özay Ezerceli, Reyhan Bayraktar, Fatma Betül Terzioğlu

机构 * Hidden for Review(保密)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16696 2025-09-23 cs.CL cs.LG 62%

Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models

Wataru Hashimoto, Hidetaka Kamigaito, Taro Watanabe

机构 * Nara Institute of Science and Technology(奈良科学技术大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted at EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13702 2025-09-18 cs.CL cs.AI 62%

DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models

Xiao Zheng

机构 * School of Computing and Technology(计算机学院) China University of Petroleum(中国石油大学) Qingdao(青岛)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13334 2025-09-18 cs.AI cs.LG 62%

FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness

Anand Swaroop, Akshat Nallani, Saksham Uboweja, Adiliia Uzdenova, Michael Nguyen, Kevin Zhu, Sunishchal Dev, Ashwinee Panda, Vasu Sharma, Maheep Chaudhary

机构 * Algoverse AI Research(Algoverse AI研究院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07082 2025-09-16 cs.CV cs.AI cs.LG 62%

On the Generalization of Representation Uncertainty in Earth Observation

Spyros Kondylatos, Nikolaos Ioannis Bountos, Dimitrios Michail, Xiao Xiang Zhu, Gustau Camps-Valls, Ioannis Papoutsis

机构 * National Observatory of Athens(雅典国家天文台) National Technical University of Athens(雅典技术大学) University of Valencia(瓦伦西亚大学) Harokopio University of Athens(雅典惠克罗波利斯大学) Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Archimedes/Athena RC(阿基米德/雅典RC)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01222 2025-09-16 cs.LG cs.AI 62%

Calibration in Deep Learning: A Survey of the State-of-the-Art

Cheng Wang

机构 * Amazon(亚马逊)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10208 2025-09-15 cs.CL cs.AI 62%

SI-FACT: Mitigating Knowledge Conflict via Self-Improving Faithfulness-Aware Contrastive Tuning

Shengqiang Fu

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15850 2025-09-12 cs.LG cs.AI 62%

Uncertainty Estimation by Human Perception versus Neural Models

Pedro Mendes, Paolo Romano, David Garlan

机构 * Software and Societal Systems Department, Carnegie Mellon University(卡内基梅隆大学软件与社会系统部门) INESC-ID and Instituto Superior Técnico, Universidade de Lisboa(里斯本大学INESC-ID和理工学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00300 2025-09-11 cs.HC cs.AI cs.LG 62%

MetaExplainer: A Framework to Generate Multi-Type User-Centered Explanations for AI Systems

Shruthi Chari, Oshani Seneviratne, Prithwish Chakraborty, Pablo Meyer, Deborah L. McGuinness

机构 * Rensselaer Polytechnic Institute(伦斯勒理工学院) Amazon Science(亚马逊科学) IBM Research(IBM研究院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07475 2025-09-10 cs.CL cs.AI 62%

HALT-RAG: A Task-Adaptable Framework for Hallucination Detection with Calibrated NLI Ensembles and Abstention

Saumya Goswami, Siddharth Kurra

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06596 2025-09-09 cs.CL cs.AI 62%

HAVE: Head-Adaptive Gating and ValuE Calibration for Hallucination Mitigation in Large Language Models

Xin Tong, Zhi Lin, Jingya Wang, Bo Jin

机构 * Xin Tong School of Information and Cyber Security People’s Public Security University of China(信息与网络安全学院 中国人民公安大学) Zhi Lin School of Safety Science Tsinghua University(安全科学学院 清华大学) Jingya Wang School of Information and Cyber Security People’s Public Security University of China(信息与网络安全学院 中国人民公安大学) Bo Jin* The Third Research Institute of the Ministry of Public Security of China(公安部第三研究所)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏