arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1755 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1755 篇

2501.03282 2025-01-08 cs.AI cs.LG 62%

From Aleatoric to Epistemic: Exploring Uncertainty Quantification Techniques in Artificial Intelligence

Tianyang Wang, Yunze Wang, Jun Zhou, Benji Peng, Xinyuan Song, Charles Zhang, Xintian Sun, Qian Niu, Junyu Liu, Silin Chen, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Ming Liu, Yichao Zhang, Cheng Fei, Caitlyn Heqi Yin, Lawrence KQ Yan

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10236 2025-01-07 cs.SE cs.AI cs.CL 62%

Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models

Yuheng Huang, Jiayang Song, Zhijie Wang, Shengming Zhao, Huaming Chen, Felix Juefei-Xu, Lei Ma

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

Comments Update website, code, and experiments on eight new LLMs. To appear in the IEEE Transactions on Software Engineering (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18672 2024-12-30 cs.CL cs.AI 62%

From Hallucinations to Facts: Enhancing Language Models with Curated Knowledge Graphs

Ratnesh Kumar Joshi, Sagnik Sengupta, Asif Ekbal

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 Pages, 5 Tables, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07213 2024-12-24 cs.HC cs.CL cs.CY 62%

Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AI

Houjiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou, Daisy Pinaroc, Matthew Lease, Min Kyung Lee

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.CY

Comments Accepted at CSCW 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12767 2024-12-18 cs.AI cs.CL 62%

A Survey of Calibration Process for Black-Box LLMs

Liangru Xie, Hui Liu, Jingying Zeng, Xianfeng Tang, Yan Han, Chen Luo, Jing Huang, Zhen Li, Suhang Wang, Qi He

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07255 2024-12-11 cs.CL cs.AI 62%

Label-Confidence-Aware Uncertainty Estimation in Natural Language Generation

Qinhong Lin, Linna Zhou, Zhongliang Yang, Yuang Cai

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16489 2024-11-26 cs.CL cs.AI 62%

O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Zhen Huang, Haoyang Zou, Xuefeng Li, Yixiu Liu, Yuxiang Zheng, Ethan Chern, Shijie Xia, Yiwei Qin, Weizhe Yuan, Pengfei Liu

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13845 2024-11-04 cs.CL cs.AI 62%

Semantic Density: Uncertainty Quantification for Large Language Models through Confidence Measurement in Semantic Space

Xin Qiu, Risto Miikkulainen

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to Neurips 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09335 2024-10-28 cs.CL cs.AI 62%

Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization

George Chrysostomou, Zhixue Zhao, Miles Williams, Nikolaos Aletras

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

Comments TACL 2024 (Presented at EMNLP 2024)

Journal ref Transactions of the Association for Computational Linguistics (2024) 12: 1163-1181

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17819 2024-10-21 cs.CL cs.AI 62%

Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation

Magdalena Wysocka, Oskar Wysocki, Maxime Delmas, Vincent Mutel, Andre Freitas

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at the Journal of Biomedical Informatics, Volume 158, October 2024, 104724

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10392 2024-10-15 cs.AI cs.CL 62%

Optimizing Instruction Synthesis: Effective Exploration of Evolutionary Space with Tree Search

Chenglin Li, Qianglong Chen, Zhi Li, Feng Tao, Yicheng Li, Hao Chen, Fei Yu, Yin Zhang

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02773 2024-10-07 cs.CV cs.AI cs.CL 62%

Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA

Jian Lan, Diego Frassinelli, Barbara Plank

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07950 2024-10-04 cs.CL cs.AI cs.HC 62%

Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance

Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, Nouha Dziri, Dan Jurafsky, Maarten Sap

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17405 2024-09-27 cs.AI cs.LG 62%

AI Enabled Neutron Flux Measurement and Virtual Calibration in Boiling Water Reactors

Anirudh Tunga, Jordan Heim, Michael Mueterthies, Thomas Gruenwald, Jonathan Nistor

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Journal ref 13th Nuclear Plant Instrumentation, Control & Human-Machine Interface Technologies (NPIC&HMIT 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13585 2024-09-23 cs.LG cs.AI 62%

Neurosymbolic Conformal Classification

Arthur Ledaguenel, Céline Hudelot, Mostepha Khouadjia

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 0 figures. arXiv admin note: text overlap with arXiv:2404.08404

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09056 2024-09-11 cs.AI cs.LG 62%

Is Epistemic Uncertainty Faithfully Represented by Evidential Deep Learning Methods?

Mira Jürgens, Nis Meinert, Viktor Bengs, Eyke Hüllermeier, Willem Waegeman

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref Proceedings of the 41st International Conference on Machine Learning (ICML), 2024, pp. 22624--22642

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06736 2024-08-14 cs.CY cs.AI 62%

Speculations on Uncertainty and Humane Algorithms

Nicholas Gray

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02706 2024-08-07 cs.LG cs.AI 62%

Bayesian Kolmogorov Arnold Networks (Bayesian_KANs): A Probabilistic Approach to Enhance Accuracy and Interpretability

Masoud Muhammed Hassan

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01301 2024-08-05 stat.ML cs.AI cs.LG 62%

A Decision-driven Methodology for Designing Uncertainty-aware AI Self-Assessment

Gregory Canal, Vladimir Leung, Philip Sage, Eric Heim, I-Jeng Wang

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01168 2024-08-05 cs.CL cs.AI 62%

Misinforming LLMs: vulnerabilities, challenges and opportunities

Bo Zhou, Daniel Geißler, Paul Lukowicz

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00550 2024-08-02 cs.CV cs.AI cs.CL 62%

Mitigating Multilingual Hallucination in Large Vision-Language Models

Xiaoye Qu, Mingyang Song, Wei Wei, Jianfeng Dong, Yu Cheng

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18370 2024-07-29 cs.LG cs.CL 62%

Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

Jaehun Jung, Faeze Brahman, Yejin Choi

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12330 2024-07-18 cs.LG cs.AI 62%

Uncertainty Calibration with Energy Based Instance-wise Scaling in the Wild Dataset

Mijoo Kim, Junseok Kwon

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02662 2024-07-04 cs.SI cs.CL cs.CY 62%

Supporters and Skeptics: LLM-based Analysis of Engagement with Mental Health (Mis)Information Content on Video-sharing Platforms

Viet Cuong Nguyen, Mini Jain, Abhijat Chauhan, Heather Jaime Soled, Santiago Alvarez Lesmes, Zihang Li, Michael L. Birnbaum, Sunny X. Tang, Srijan Kumar, Munmun De Choudhury

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.CY

Comments 12 pages, in submission to ICWSM

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.21028 2024-07-04 cs.CL cs.AI 62%

LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models

Elias Stengel-Eskin, Peter Hase, Mohit Bansal

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 18 pages. Code: https://github.com/esteng/pragmatic_calibration

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01122 2024-07-02 cs.CL cs.LG 62%

Calibrated Large Language Models for Binary Question Answering

Patrizio Giovannotti, Alexander Gammerman

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted to COPA 2024 (13th Symposium on Conformal and Probabilistic Prediction with Applications)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11353 2024-06-18 cs.LG cs.CL 62%

$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts

Guanjie Chen, Xinyu Zhao, Tianlong Chen, Yu Cheng

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.LG

Comments 9 pages, 8 figures, camera ready on ICML2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12350 2024-06-14 cs.AI cs.CY econ.GN q-fin.EC 62%

Artificial Intelligence and Dual Contract

Qian Qi

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18715 2024-06-06 cs.CV cs.AI cs.CL cs.MM 62%

Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding

Xintong Wang, Jingheng Pan, Liang Ding, Chris Biemann

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to Findings of ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13006 2024-06-05 cs.LG cs.CL 62%

Investigating the Impact of Model Instability on Explanations and Uncertainty

Sara Vera Marjanović, Isabelle Augenstein, Christina Lioma

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏