arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1753 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1753 篇

2310.03525 2025-09-03 cs.CV 67%

Vehicle-to-Everything Cooperative Perception for Autonomous Driving

Tao Huang, Jianan Liu, Xi Zhou, Dinh C. Nguyen, Mostafa Rahimi Azghadi, Yuxuan Xia, Qing-Long Han, Sumei Sun

机构 * College of Science and Engineering, James Cook University(科学与工程学院,詹姆斯库克大学) Momoni AI Department of Electrical and Computer Engineering, University of Alabama in Huntsville(电气与计算机工程系,阿拉巴马大学亨茨维尔分校) Department of Automation and Intelligent Sensing, Shanghai Jiaotong University(自动化与智能感知系,上海交通大学) School of Engineering, Swinburne University of Technology(工程学院,斯威本技术大学) Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR)(信息与通信研究机构,科技研究局(A*STAR))

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract)

Comments This article has been accepted for publication in Proceedings of the IEEE on 11 August 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02319 2025-08-27 cs.CV 67%

Is Uncertainty Quantification a Viable Alternative to Learned Deferral?

Anna M. Wundram, Christian F. Baumgartner

机构 * Faculty of Health Sciences and Medicine, University of Lucerne, Switzerland(健康科学与医学学院,卢塞恩大学,瑞士) Cluster of Excellence -- ML for Science, University of Tübingen, Germany(卓越中心——科学中的机器学习,图宾根大学,德国)

专题命中 幻觉与事实性 :safety(abstract);AI safety(abstract)

Comments Accepted as an oral presentation at MICCAI UNSURE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17419 2025-06-24 cs.CL cs.AI cs.LG stat.ML 67%

UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making

Jinhao Duan, James Diffenderfer, Sandeep Madireddy, Tianlong Chen, Bhavya Kailkhura, Kaidi Xu

机构 * Drexel University(德雷塞尔大学) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) Argonne National Laboratory(阿贡国家实验室) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 19 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11001 2025-06-16 cs.SE cs.AI cs.CY cs.LG 67%

Rethinking Technological Readiness in the Era of AI Uncertainty

S. Tucker Browne, Mark M. Bailey

机构 * United States Air Force(美国空军) National Intelligence University(国家情报大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17004 2025-06-03 cs.LG cs.AI cs.CL stat.ML 67%

(Im)possibility of Automated Hallucination Detection in Large Language Models

Amin Karbasi, Omar Montasser, John Sous, Grigoris Velegkas

机构 * Yale University(耶鲁大学)

专题命中 幻觉与事实性 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05980 2025-06-03 cs.CL cs.AI cs.LG 67%

FactLens: Benchmarking Fine-Grained Fact Verification

Kushan Mitra, Dan Zhang, Sajjadur Rahman, Estevam Hruschka

机构 * Megagon Labs(梅加贡实验室) Adobe Inc.(Adobe公司)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 12 pages, updated version

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23968 2025-06-02 cs.CR cs.AI cs.CY cs.LG stat.ML 67%

Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention

Stephan Rabanser, Ali Shahin Shamsabadi, Olive Franzese, Xiao Wang, Adrian Weller, Nicolas Papernot

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Proceedings of the 42nd International Conference on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23854 2025-06-02 cs.CL cs.AI cs.LG 67%

Revisiting Uncertainty Estimation and Calibration of Large Language Models

Linwei Tao, Yi-Fan Yeh, Minjing Dong, Tao Huang, Philip Torr, Chang Xu

机构 * School of Computer Science University of Sydney(悉尼大学计算机科学学院) City University of Hong Kong(香港城市大学) Shanghai Jiao Tong University(上海交通大学) Department of Engineering Science University of Oxford(牛津大学工程科学系)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22655 2025-05-29 cs.LG cs.AI cs.CL 67%

Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents

Michael Kirchhof, Gjergji Kasneci, Enkelejda Kasneci

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17671 2025-05-16 cs.CL cs.AI cs.LG 67%

Data-Driven Calibration of Prediction Sets in Large Vision-Language Models Based on Inductive Conformal Prediction

Yuanchang Ye, Weiyan Wen

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by ICIPCA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06529 2025-02-12 cs.AI cs.CL cs.LG 67%

Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity

Kaiqu Liang, Zixu Zhang, Jaime Fernández Fisac

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04428 2025-02-10 cs.CL cs.AI cs.LG 67%

Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization

Yu-Neng Chuang, Leisheng Yu, Guanchu Wang, Lizhe Zhang, Zirui Liu, Xuanting Cai, Yang Sui, Vladimir Braverman, Xia Hu

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12215 2025-01-28 cs.LG cs.CL cs.CY 67%

SoK: Machine Learning for Misinformation Detection

Madelyne Xiao, Jonathan Mayer

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07961 2024-12-12 cs.CL cs.AI cs.LG 67%

Forking Paths in Neural Text Generation

Eric Bigelow, Ari Holtzman, Hidenori Tanaka, Tomer Ullman

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00760 2024-12-03 eess.AS cs.AI cs.CL cs.ET cs.LG 67%

Automating Feedback Analysis in Surgical Training: Detection, Categorization, and Assessment

Firdavs Nasriddinov, Rafal Kocielnik, Arushi Gupta, Cherine Yang, Elyssa Wong, Anima Anandkumar, Andrew Hung

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted as a proceedings paper at Machine Learning for Health 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23054 2024-11-25 cs.LG cs.AI cs.CL cs.CV 67%

Controlling Language and Diffusion Models by Transporting Activations

Pau Rodriguez, Arno Blaas, Michal Klein, Luca Zappella, Nicholas Apostoloff, Marco Cuturi, Xavier Suau

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00499 2024-11-19 cs.CL cs.AI cs.LG 67%

ConU: Conformal Uncertainty in Large Language Models with Correctness Coverage Guarantees

Zhiyuan Wang, Jinhao Duan, Lu Cheng, Yue Zhang, Qingni Wang, Xiaoshuang Shi, Kaidi Xu, Hengtao Shen, Xiaofeng Zhu

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14259 2024-11-19 cs.CL cs.AI cs.LG 67%

Word-Sequence Entropy: Towards Uncertainty Estimation in Free-Form Medical Question Answering Applications and Beyond

Zhiyuan Wang, Jinhao Duan, Chenxi Yuan, Qingyu Chen, Tianlong Chen, Yue Zhang, Ren Wang, Xiaoshuang Shi, Kaidi Xu

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by Engineering Applications of Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04331 2024-11-06 cs.CL cs.AI cs.IR cs.LG 67%

PaCE: Parsimonious Concept Engineering for Large Language Models

Jinqi Luo, Tianjiao Ding, Kwan Ho Ryan Chan, Darshan Thaker, Aditya Chattopadhyay, Chris Callison-Burch, René Vidal

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted in NeurIPS 2024. GitHub repository at https://github.com/peterljq/Parsimonious-Concept-Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02062 2024-10-28 cs.CL cs.AI cs.LG 67%

Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?

Wataru Hashimoto, Hidetaka Kamigaito, Taro Watanabe

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2024 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16592 2024-10-23 cs.LG cs.CL cs.CY 67%

ViMGuard: A Novel Multi-Modal System for Video Misinformation Guarding

Andrew Kan, Christopher Kan, Zaid Nabulsi

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.CY、cs.LG

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10474 2024-08-21 cs.SE cs.AI cs.CL cs.CR cs.LG 67%

LeCov: Multi-level Testing Criteria for Large Language Models

Xuan Xie, Jiayang Song, Yuheng Huang, Da Song, Fuyuan Zhang, Felix Juefei-Xu, Lei Ma

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07453 2024-08-15 cs.CL cs.AI cs.LG 67%

Fact or Fiction? Improving Fact Verification with Knowledge Graphs through Simplified Subgraph Retrievals

Tobias A. Opsahl

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 10 pages, 3 figures, appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12836 2024-07-19 cs.CL cs.AI cs.LG 67%

OSPC: Artificial VLM Features for Hateful Meme Detection

Peter Grönquist

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00030 2024-06-04 cs.CL cs.AI cs.LG 67%

Large Language Model Pruning

Hanjuan Huang, Hao-Jia Song, Hsing-Kuo Pao

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 17 pages, 7 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20003 2024-05-31 cs.LG cs.AI cs.CL 67%

Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities

Alexander Nikitin, Jannik Kossen, Yarin Gal, Pekka Marttinen

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13788 2024-04-25 cs.CL cs.AI cs.CR cs.HC cs.LG 67%

Can LLM-Generated Misinformation Be Detected?

Canyu Chen, Kai Shu

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to Proceedings of ICLR 2024. 9 pages for main paper, 40 pages including appendix. The code, results, dataset for this paper and more resources on "LLMs Meet Misinformation" have been released on the project website: https://llm-misinformation.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08705 2024-04-16 cs.CL cs.AI cs.LG 67%

Introducing L2M3, A Multilingual Medical Large Language Model to Advance Health Equity in Low-Resource Regions

Agasthya Gangavarapu

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05973 2024-03-12 cs.CL cs.AI cs.LG 67%

Calibrating Large Language Models Using Their Generations Only

Dennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun, Seong Joon Oh

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03170 2024-03-12 cs.MM cs.AI cs.CL cs.CV cs.CY 67%

SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection

Peng Qi, Zehong Yan, Wynne Hsu, Mong Li Lee

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments To appear in CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏