arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1755 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1755 篇

2509.18792 2025-09-24 cs.CL 57%

Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing

Sabri Boughorbel, Fahim Dalvi, Nadir Durrani, Majd Hawasly

机构 * Qatar Computing Research Institute, HBKU(卡塔尔计算研究所,哈瓦那大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments 12 pages, accepted to the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18132 2025-09-24 cs.AI 57%

Position Paper: Integrating Explainability and Uncertainty Estimation in Medical AI

Xiuyi Fan

机构 * Lee Kong Chian School of Medicine, College of Computing Data Science, Nanyang Technological University, Singapore

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments Accepted at the International Joint Conference on Neural Networks, IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18128 2025-09-24 cs.LG 57%

Accounting for Uncertainty in Machine Learning Surrogates: A Gauss-Hermite Quadrature Approach to Reliability Analysis

Amirreza Tootchi, Xiaoping Du

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16742 2025-09-23 cs.AI 57%

Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories

Mohammad Beigi, Ying Shen, Parshin Shojaee, Qifan Wang, Zichao Wang, Chandan Reddy, Ming Jin, Lifu Huang

机构 * University of California, Davis(加州大学戴维斯分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Virginia Tech(弗吉尼亚理工大学) Meta AI Adobe Research(Adobe研究)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16369 2025-09-23 cs.IR cs.AI cs.CE 57%

Enhancing Financial RAG with Agentic AI and Multi-HyDE: A Novel Approach to Knowledge Retrieval and Hallucination Reduction

Akshay Govind Srinivasan, Ryan Jacob George, Jayden Koshy Joe, Hrushikesh Kant, Harshith M R, Sachin Sundar, Sudharshan Suresh, Rahul Vimalkanth, Vijayavallabh

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments 14 Pages, 8 Tables, 2 Figures. Accepted and to be published in the proceedings of FinNLP, Empirical Methods in Natural Language Processing 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07755 2025-09-23 cs.CL cs.CR 57%

Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts

Rochana Prih Hastuti, Rian Adam Rajagede, Mansour Al Ghanim, Mengxin Zheng, Qian Lou

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments Accepted at EMNLP 2025 Findings. Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12406 2025-09-17 cs.LG stat.ML 57%

Bayesian Parametric Matrix Models: Principled Uncertainty Quantification for Spectral Learning

Mohammad Nooraiepour

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12034 2025-09-16 cs.AI 57%

Human-AI Use Patterns for Decision-Making in Disaster Scenarios: A Systematic Review

Emmanuel Adjei Domfeh, Christopher L. Dancy

机构 * Department of Computer Science and Engineering(计算机科学与工程系) The Pennsylvania State University(宾夕法尼亚州立大学) Department of Industrial and Manufacturing Engineering(工业与制造工程系)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08777 2025-09-11 cs.CV cs.CL 57%

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles

Eric Slyman, Mehrab Tanjim, Kushal Kafle, Stefan Lee

机构 * Adobe(Adobe公司) Oregon State University(俄勒冈州立大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments 17 pages, 8 figures, Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08679 2025-09-11 cs.LG 57%

Signal Fidelity Index-Aware Calibration for Dementia Predictions Across Heterogeneous Real-World Data

Jingya Cheng, Jiazi Tian, Federica Spoto, Alaleh Azhir, Daniel Mork, Hossein Estiri

机构 * Department of Medicine, Massachusetts General Hospital(麻省总医院内科部) Department of Biostatistics, Harvard T.H. Chan School of Public Health(哈佛T.H. Chan公共卫生学院生物统计学部) Department of Medicine, Brigham and Women’s Hospital(布里洛妇产科医院内科部)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13107 2025-09-11 cs.CL cs.IR 57%

All for law and law for all: Adaptive RAG Pipeline for Legal Research

Figarri Keisha, Prince Singh, Pallavi, Dion Fernandes, Aravindh Manivannan, Ilham Wicaksono, Faisal Ahmad, Wiem Ben Rim

机构 * University College London(伦敦大学学院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments submitted to NLLP 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05557 2025-09-09 cs.AI 57%

MV-Debate: Multi-view Agent Debate with Dynamic Reflection Gating for Multimodal Harmful Content Detection in Social Media

Rui Lu, Jinhe Bi, Yunpu Ma, Feng Xiao, Yuntao Du, Yijun Tian

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04735 2025-09-08 cs.CV cs.AI 57%

Enhancing Self-Driving Segmentation in Adverse Weather Conditions: A Dual Uncertainty-Aware Training Approach to SAM Optimization

Dharsan Ravindran, Kevin Wang, Zhuoyuan Cao, Saleh Abdelrahman, Jeffery Wu

机构 * Queen's University(女王大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04664 2025-09-08 cs.CL 57%

Why Language Models Hallucinate

Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang

机构 * OpenAI Georgia Tech(佐治亚理工学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02129 2025-09-03 cs.LG cs.CV 57%

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time

Jintao Cheng, Weibin Li, Jiehao Luo, Xiaoyu Tang, Zhijian He, Jin Wu, Yao Zou, Wei Zhang

机构 * Hong Kong University of Science(香港科学与技术大学) South China Normal University, Shanwei, Guangdong, China(华南师范大学,汕尾,广东,中国) Shenzhen Technology University, Shenzhen, Guangdong, China(深圳科技大学,深圳,广东,中国) University of Science(科学大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00869 2025-09-03 cs.CL 57%

Exploring and Mitigating Fawning Hallucinations in Large Language Models

Zixuan Shangguan, Yanjie Dong, Lanjun Wang, Xiaoyi Fan, Victor C. M. Leung, Xiping Hu

机构 * School of Medical Technology, Beijing Institute of Technology(北京理工大学医学技术学院) Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(深圳MSU-BIT大学人工智能研究院) Tianjin University(天津大学) The Hong Kong University of Science(香港科学大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11649 2025-09-03 cs.AI cs.SI 57%

Competing LLM Agents in a Non-Cooperative Game of Opinion Polarisation

Amin Qasmi, Usman Naseem, Mehwish Nasim

机构 * The University of Western Australia(西澳大学) Lahore University of Management Sciences(拉合尔管理科学大学) Macquarie University(麦考瑞大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12067 2025-09-01 eess.AS cs.AI cs.SD 57%

Evaluating Logit-Based GOP Scores for Mispronunciation Detection

Aditya Kamlesh Parikh, Cristian Tejedor-Garcia, Catia Cucchiarini, Helmer Strik

机构 * Centre for Language Studies(语言研究学院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

Comments Accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)

Journal ref https://www.isca-archive.org/interspeech_2025/parikh25b_interspeech.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19432 2025-08-28 cs.AI 57%

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs

Yao Fu, Xianxuan Long, Runchao Li, Haotian Yu, Mu Sheng, Xiaotian Han, Yu Yin, Pan Li

机构 * Case Western Reserve University(凯斯西储大学) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

Comments Accepted to EMNLP2025 main conference (poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12964 2025-08-26 cs.CL 57%

Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Adi Simhi, Itay Itzhak, Fazl Barez, Gabriel Stanovsky, Yonatan Belinkov

机构 * Technion – Israel Institute of Technology(技术ion-以色列理工学院) University of Oxford and WhiteBox(牛津大学和WhiteBox) School of Computer Science and Engineering, The Hebrew University of Jerusalem(耶路撒冷希伯来大学计算机科学与工程学院)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18001 2025-08-26 cs.LG stat.ML 57%

A Novel Framework for Uncertainty Quantification via Proper Scores for Classification and Beyond

Sebastian G. Gruber

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments PhD Thesis (cumulative, spanning 6 peer-reviewed publications)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01225 2025-08-25 cs.CV cs.AI 57%

Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language Models

Xinyu Chen, Haotian Zhai, Can Zhang, Xiupeng Shi, Ruirui Li

机构 * Shanghai University(上海大学) Beijing University of Chemical Technology(北京化工大学) University of Minnesota(明尼苏达大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14266 2025-08-21 cs.CV cs.AI 57%

Effect of Data Augmentation on Conformal Prediction for Diabetic Retinopathy

Rizwan Ahamed, Annahita Amireskandari, Joel Palko, Carol Laxson, Binod Bhattarai, Prashnna Gyawali

机构 * West Virginia University(西弗吉尼亚大学) University of Aberdeen(阿伯丁大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments 3rd Workshop in Data Engineering in Medical Imaging (DEMI), MICCAI-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11257 2025-08-18 cs.SE cs.AI 57%

Hallucination in LLM-Based Code Generation: An Automotive Case Study

Marc Pavel, Nenad Petrovic, Lukasz Mazur, Vahid Zolfaghari, Fengjunjie Pan, Alois Knoll

机构 * Real-Time Systems Technical University of Munich(实时系统技术大学慕尼黑)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10010 2025-08-15 cs.CL 57%

An Audit and Analysis of LLM-Assisted Health Misinformation Jailbreaks Against LLMs

Ayana Hussain, Patrick Zhao, Nicholas Vincent

专题命中 幻觉与事实性 :jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09458 2025-08-15 cs.HC cs.AI cs.ET 57%

Hallucination vs interpretation: rethinking accuracy and precision in AI-assisted data extraction for knowledge synthesis

Xi Long, Christy Boscardin, Lauren A. Maggio, Joseph A. Costello, Ralph Gonzales, Rasmyah Hammoudeh, Ki Lai, Yoon Soo Park, Brian C. Gin

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09561 2025-08-14 cs.LG 57%

Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges

Changyuan Zhao, Guangyuan Liu, Ruichen Zhang, Yinqiu Liu, Jiacheng Wang, Jiawen Kang, Dusit Niyato, Zan Li, Xuemin, Shen, Zhu Han, Sumei Sun, Chau Yuen, Dong In Kim

机构 * College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) School of Automation, Guangdong University of Technology(广东工业大学自动化学院) State Key Laboratory of Integrated Services Networks, Xidian University(西安电子科技大学集成服务网络国家重点实验室) Department of Electrical and Computer Engineering, University of Waterloo(滑铁卢大学电气与计算机工程系) Department of Computer Science and Engineering, Kyung Hee University(韩国庆熙大学计算机科学与工程系) Institute for Infocomm Research, Agency for Science, Technology and Research(科技研究局信息通信研究所) Department of Electrical and Computer Engineering, Sungkyunkwan University(庆熙大学电气与计算机工程系)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

Comments 21 pages. 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06795 2025-08-13 cs.CL cs.CV 57%

From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models

Yuying Shang, Xinyi Zeng, Yutao Zhu, Xiao Yang, Zhengwei Fang, Jingyuan Zhang, Jiawei Chen, Zinan Liu, Yu Tian

机构 * University of Chinese Academy of Sciences(中国科学院大学) Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University(计算机科学与技术系,人工智能研究院,清华大学) Gaoling School of Artificial Intelligence, Renmin University of China(人工智能学院,中国人民大学) Kuaishou Technology Inc.(快手科技有限公司) Shanghai Key Laboratory of Multi. Info. Processing, East China Normal University(多信息处理重点实验室,华东师范大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07223 2025-08-12 cs.IR cs.AI 57%

Selection and Exploitation of High-Quality Knowledge from Large Language Models for Recommendation

Guanchen Wang, Mingming Ha, Tianbao Ma, Linxun Chen, Zhaojie Liu, Guorui Zhou, Kun Gai

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04105 2025-08-07 cs.AI 57%

Towards Transparent AI Grading: Semantic Entropy as a Signal for Human-AI Disagreement

Karrtik Iyer, Manikandan Ravikiran, Prasanna Pendse, Shayan Mohanty

机构 * Thoughtworks AI Research Labs(Thoughtworks人工智能研究实验室)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏