arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1753 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1753 篇

2311.09114 2024-02-27 cs.CL cs.AI cs.LG 67%

Ever: Mitigating Hallucination in Large Language Models through Real-Time Verification and Rectification

Haoqiang Kang, Juntong Ni, Huaxiu Yao

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09731 2024-02-19 cs.CL cs.AI cs.LG 67%

Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge

Genglin Liu, Xingyao Wang, Lifan Yuan, Yangyi Chen, Hao Peng

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07023 2024-02-13 cs.CL cs.AI cs.CV cs.HC cs.LG 67%

Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations

Ankit Pal, Malaikannan Sankarasubbu

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Preprint version, Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04782 2023-12-05 cs.CL cs.AI cs.LG 67%

HistAlign: Improving Context Dependency in Language Generation by Aligning with History

David Wan, Shiyue Zhang, Mohit Bansal

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2023 (20 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08401 2023-11-15 cs.CL cs.AI cs.LG 67%

Fine-tuning Language Models for Factuality

Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning, Chelsea Finn

专题命中 幻觉与事实性 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10346 2023-09-20 cs.LG cs.AI cs.CL 67%

Explaining Agent Behavior with Large Language Models

Xijia Zhang, Yue Guo, Simon Stepputtis, Katia Sycara, Joseph Campbell

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Human Multi-Robot Interaction Workshop at IROS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01906 2023-08-04 cs.CL cs.AI cs.LG 67%

Reasoning in Large Language Models Through Symbolic Math Word Problems

Vedant Gaur, Nikunj Saunshi

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at the Findings of ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15703 2023-07-31 cs.CL cs.AI cs.LG 67%

Uncertainty in Natural Language Generation: From Theory to Applications

Joris Baan, Nico Daheim, Evgenia Ilia, Dennis Ulmer, Haau-Sing Li, Raquel Fernández, Barbara Plank, Rico Sennrich, Chrysoula Zerva, Wilker Aziz

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.11768 2023-07-26 cs.CL cs.AI cs.LG 67%

Question Decomposition Improves the Faithfulness of Model-Generated Reasoning

Ansh Radhakrishnan, Karina Nguyen, Anna Chen, Carol Chen, Carson Denison, Danny Hernandez, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Sam McCandlish, Sheer El Showk, Tamera Lanham, Tim Maxwell, Venkatesa Chandrasekaran, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, Ethan Perez

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments For few-shot examples and prompts, see https://github.com/anthropics/DecompositionFaithfulnessPaper

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08216 2022-12-20 cs.LG cs.AI cs.CL cs.HC 67%

Azimuth: Systematic Error Analysis for Text Classification

Gabrielle Gauthier-Melançon, Orlando Marquez Ayala, Lindsay Brin, Chris Tyler, Frédéric Branchaud-Charron, Joseph Marinier, Karine Grande, Di Le

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments To be published in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 13 pages and 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04272 2022-12-09 cs.LG cs.AI cs.CL cs.SI 67%

A Modality-level Explainable Framework for Misinformation Checking in Social Networks

Vítor Lourenço, Aline Paes

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to publication at LatinX in AI workshop at the Thirty-sixth Conference on Neural Information Processing Systems, LXAI @ NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.08884 2020-12-21 cs.CL cs.AI cs.LG 67%

Learning from the Best: Rationalizing Prediction by Adversarial Information Calibration

Lei Sha, Oana-Maria Camburu, Thomas Lukasiewicz

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Journal ref Proceedings of the 35th AAAI Conference on Artificial Intelligence, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03296 2026-05-21 cs.CL cs.AI 66%

A Systematic Comparison between Extractive Self-Explanations and Human Rationales in Text Classification

抽取式自我解释与人类推理在文本分类中的系统比较

Stephanie Brandl, Oliver Eberle

机构 * Center for Social Data Science(社会科学数据科学中心) University of Copenhagen(哥本哈根大学) Machine Learning Group(机器学习小组) Technische Universität Berlin(柏林技术大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI;trustworthy(comments)

AI总结 本文比较了抽取式自我解释与人类推理在文本分类任务中的有效性,通过分析不同任务和语言的解释质量,发现自我解释在文本长度和任务复杂度上与人类推理存在显著差异。

Comments accepted to the Trustworthy NLP Workshop, co-located with ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10573 2026-08-24 cs.CL cs.LG 版本更新 62%

The Intrinsic Dimension of Prompts in Internal Representations of Large Language Models

大语言模型内部表示中提示的固有维度

Karthik Viswanathan, Yuri Gardinazzi, Giada Panerai, Alberto Cazzaniga, Matteo Biagetti

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.LG

AI总结 本研究通过固有维度视角分析大语言模型提示的内部表示几何,发现其与下一个 token 不确定性等相关,训练的线性探测模型可在生成前区分恶意与良性提示,准确率优于现有安全工具,为 LLM 安全提供了新信号。

Comments 12+14 pages, 18 figures, matches published version on Transactions of Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19206 2026-08-21 cs.CL cs.AI cs.MA 新提交 62%

Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses

幻觉作为特征而非缺陷:评估一种将推测性语言模型输出转化为可检验科学假说的多智能体架构

Nicolas Rodriguez-Alvarez

机构 * IES Parquesol(帕尔奎索尔学院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 该研究提出基于Rust的多智能体架构,通过生成与评估智能体的认识论摩擦循环将LLM的推测性输出转化为可检验假说,实验显示该架构在需经受严格约束时更具优势,且各架构在原创性等维度的平衡表现不同。

Comments 25 pages. Bilingual: full English version followed by the complete Spanish version. Includes an exploratory paired baseline and ablation study (6 conditions). Code and data: https://doi.org/10.5281/zenodo.20649714

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11922 2026-08-21 cs.CL cs.IR cs.LG 版本更新 62%

LODESTAR: Robust Entropy-Based Answer Selection in Retrieval-Augmented Generation for Question Answering -- Directing Frozen-LLM Entropy with a Reinforcement-Learned Prompt Polarizer under Misleading Passages

LODESTAR:可信赖熵是被引导的,而非仅被测量——强化偏振器防止冻结型大语言模型(LLM)被错误证据误导而自信出错

Hung-Chun Hsu, Po-Jen Ko, Che-Cheng Wu, Li-Yang Chang, Chuan-Ju Wang

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.LG

AI总结 LODESTAR是首个通过强化学习训练偏振器、依据第三方冻结型LLM的不确定性评分文本干预的方法,可提升检索增强问答的F1值等指标,减少模型读取误导性段落的概率。

Comments 28 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18709 2026-08-20 cs.CV cs.AI cs.LG 新提交 62%

A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation

语义分割中不确定性量化与基础模型的批判性综合研究

Steven Landgraf, Joceline Hinz, Markus Ulrich

机构 * Institute of Photogrammetry and Remote Sensing (IPF), Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院摄影测量与遥感研究所)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文首次系统评估四类不确定性量化方法在语义分割基础模型上的表现,揭示预测性能、可靠性与计算成本间的权衡,为现实应用需联合优化相关指标指明方向。

Comments Accepted for publication in the ISPRS Annals (ISPRS Congress 2026, Toronto, Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18116 2026-08-20 cs.CL cs.LG 新提交 62%

You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language Models

你即你所提示的:农业食品视觉语言模型中的提示质量、领域偏移与不确定性

Andrea Morales-Garzón, Salvador López-Joya, Miguel López-Pérez, Maria J. Martin-Bautista

机构 * University of Granada(格拉纳达大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 该研究针对农业食品领域,评估了零样本提示集成(ZPE)在分布内与分布外场景的表现,提出PID方法提升严重领域偏移下的故障检测能力,验证了领域特定提示池的优势。

Comments Accepted in the journal Procesamiento del Lenguaje Natural (SEPLN2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16627 2026-08-18 cs.CL cs.AI 新提交 62%

When Do Explanations Help In-Context Learning? A Comparative Study of Natural Language Explanation Types and Faithfulness

解释何时能帮助上下文学习?自然语言解释类型与忠实度的比较研究

Mahdi Dhaini, Adam Dejl, Juraj Vladika, Volkan Özer, Barbara Plank, Gjergji Kasneci

机构 * LMU Munich(慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Imperial College London(伦敦帝国学院) Technical University of Munich(慕尼黑工业大学) MaiNLP lab, CIS, LMU Munich(慕尼黑大学CIS研究所MaiNLP实验室)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 该研究通过在6个基准和4个指令调优模型上的比较评估,探究不同来源与选择方式的自然语言解释对上下文学习下游性能的影响,为实际提示流程中解释的选择和报告提供了见解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15338 2026-08-18 cs.CL cs.AI 新提交 62%

When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text

当AI改写时,分类器会放松:针对讽刺文本与AI改写社交文本的不确定性感知情感分析

Shresth Shroff

机构 * Manipal University Jaipur(斋浦尔马尼帕尔大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 该研究针对讽刺与AI改写社交文本开展情感分析,发现AI改写文本可提升分类器准确率,提出轻量弃权包装器,推动高风险情感应用转向不确定性感知预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01375 2026-08-18 cs.CY cs.AI 版本更新 62%

Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics

超越访问:引导式LLM支架在本科统计学自主学习中的应用

Mohammad Amanlou, Yasaman Amou-Jafari, Fereshte Bagheri, Fatemeh Boloukazari, Mehrad Liviyan, Elahe Khodaverdi Nadrabadi, Shahab Sherafat, Behnam Bahrak

机构 * School of Electrical and Computer Engineering, University of Tehran, Iran(伊朗塔里哈大学电气与计算机工程学院) Tehran Institute for Advanced Studies, Khatam University, Iran(伊朗卡塔姆大学泰赫兰高级研究院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究通过准实验比较无LLM、无限制LLM和引导式LLM三种条件,发现引导式LLM使用能促进以推理为导向的交互模式,提升无辅助测验表现,并改善自我评估校准,表明LLM作为教育工具需通过支架设计实现推理伙伴而非答案获取工具。

Comments 10 pages. Accepted at the 34th International Conference on Computers in Education (ICCE 2026), Asia-Pacific Society for Computers in Education (APSCE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14390 2026-08-14 cs.CL cs.AI 版本更新 62%

REHEARSE: Experiential Rehearsal for Verbal Confidence Calibration in Large Language Models

REHEARSE:大语言模型中用于语言置信度校准的经验式排练

Ke Fang, Tianyi Zhao, Qianwen Wang, Lu Cheng

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Southern California(南加州大学) University of Illinois Chicago(伊利诺伊大学香槟分校)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

AI总结 针对大语言模型置信度与实际正确性不匹配的问题,提出无训练的Rehearse方法,通过置信度校准博弈的反馈生成校准信号,在多模型多基准实验中显著降低了预期校准误差并提升准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11280 2026-08-13 eess.IV cs.AI cs.CV cs.LG 新提交 62%

Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification

用于多类皮肤病变分类的不确定性感知且可解释的集成深度学习框架

Rofiqul Islam, Lilatul Ferdouse

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 5 pages, 3 figures, IEEE AIBThings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07525 2026-08-11 cs.CL cs.AI 新提交 62%

Unified Hallucination Fuzzing for Multimodal Large Language Models

面向多模态大语言模型的统一幻觉模糊测试

Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 针对多模态大语言模型的幻觉问题,提出含UniHall基准与SAMF自演化模糊测试的评估框架,发现SOTA模型在模糊测试下性能显著下降,存在有用性-幻觉权衡。

Comments 47 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07520 2026-08-11 cs.CY cs.AI 新提交 62%

KumbhDoot: A Scale-Ready, LLM-Bounded Architecture for Mass-Gathering Public-Service Assistants

KumbhDoot:面向大规模集会公共服务助手的可扩展、大语言模型(LLM)限定架构

Saurabh Sakalkar, Abhishek Singh, Ramesh Raskar

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.CY

AI总结 针对大壶节这类高风险低连接性的大规模集会,提出KumbhDoot架构,以相似度优先、LLM限定为核心,解决默认LLM助手成本高、易幻觉等问题,适配公共服务场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26502 2026-08-11 cs.AI cs.CL 版本更新 62%

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

人类放弃,推理模型坚持:将难度登记与深思分配分离

Han-yu Wang

机构 * The University of Hong Kong(香港大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文通过分离难度登记与深思分配,发现人类与大型推理模型(LRM)在问题解决中的时间/令牌分配模式相反:人类在错误问题上用时更少,而LRM在错误问题上使用更多令牌。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26798 2026-08-11 cs.LG cs.AI 版本更新 62%

Explaining, Verifying, and Aligning Semantic Hierarchies in Vision-Language Model Embeddings

解释、验证和对齐视觉语言模型嵌入中的语义层次结构

Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein, Edgar Heinert, Annika Mütze, Marvin Keller, Sparsh Tiwari, Georgii Mikriukov, Diedrich Wolter, Jae Hee Lee, Matthias Rottmann

机构 * University of Lübeck(吕贝克大学) Technical University of Munich(慕尼黑工业大学) AUMOVIO SE Osnabrück University(奥斯纳布吕克大学) Leibniz Institute for Agricultural Engineering and Bioeconomy(莱布尼茨农业工程与生物经济研究所) University of Hamburg(汉堡大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种框架,用于解释、验证和对齐视觉语言模型嵌入中的语义层次结构,通过聚类和概念库匹配提取层次结构,并利用人类本体评估其合理性,最终提出基于本体的对齐方法以提升共享嵌入空间的语义对齐。

Comments camera-ready version of accepted IJCAI-ECAI 2026 paper including supplementary material; code: https://github.com/gesina/ontological-commitment

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00151 2026-08-05 cs.CY cs.AI econ.TH 版本更新 62%

Optimising for Flourishing: Flourishing Metrics and Return on Flourishing as Success Criteria for Artificial Intelligence and Post-AGI Economic Systems

为繁荣优化:繁荣度量与繁荣回报作为人工智能及后AGI经济系统的成功标准

Keyun Ruan, Jonathan D. Teubner, John M. Bremen

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.CY

AI总结 该研究提出将人类繁荣作为AI及后AGI经济系统的核心成功标准,构建了Flourishing Metrics框架与RoF评估方法,为AI及经济转型的价值核算提供了决策架构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01012 2026-08-04 cs.CL cs.AI 新提交 62%

MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models

MedUPS:利用大语言模型辅助罕见医疗病例的诊断

Ofir Ben Shoham, Oriel Perets, Nir Grinberg, Nadav Rappoport

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出MedUPS对齐框架及MedUPSQA数据集,通过强化学习使大语言模型对齐临床中游决策,提升了不同规模模型的下一步临床决策准确率,且小模型表现优于部分大模型。

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22775 2026-08-04 cs.LG cs.AI cs.HC 版本更新 62%

MambaGaze: Bidirectional Mamba with Explicit Missing Data Modeling for Cognitive Load Assessment from Eye-Gaze Tracking Data

MambaGaze: 通过显式缺失数据建模的双向Mamba用于从眼动追踪数据中评估认知负荷

Amir Mousavi, Mohammad Sadegh Sirjani, Erfan Nourbakhsh, Mimi Xie, Rocky Slavin, Leslie Neely, John Davis, John Quarles

机构 * Department of Computer Science, College of AI, Cyber and Computing, The University of Texas at San Antonio(计算机科学系,人工智能、网络与计算学院,德克萨斯大学圣安东尼奥分校) Department of Educational Psychology, College of Education and Human Development, The University of Texas at San Antonio(教育心理学系,教育与人类发展学院,德克萨斯大学圣安东尼奥分校) Department of Neuroscience, Developmental and Regenerative Biology, College of Sciences, The University of Texas at San Antonio(神经科学系,发育与再生生物学系,科学学院,德克萨斯大学圣安东尼奥分校)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出MambaGaze,通过XMD编码和双向Mamba-2框架,解决眼动追踪数据中频繁缺失和长时序依赖建模的问题,实验证明其在认知负荷评估中的优越性能和边缘部署可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏