arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9380 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9380 篇

1912.00782 2019-12-04 cs.CY cs.AI cs.LG 82%

The relationship between trust in AI and trustworthy machine learning technologies

Ehsan Toreini, Mhairi Aitken, Kovila Coopamootoo, Karen Elliott, Carlos Gonzalez Zelaya, Aad van Moorsel

专题命中 安全评测 :trustworthy(title);safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments This submission has been accepted in ACM FAT* 2020 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05656 2026-02-10 cs.LG cs.AI 82%

Alignment Verifiability in Large Language Models: Normative Indistinguishability under Behavioral Evaluation

大语言模型对齐的可验证性:行为评估下的规范不可区分性

Igor Santos-Grueiro

机构 * Igor Santos-Grueiro(独立研究者)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了大语言模型对齐的可验证性,指出行为评估无法唯一确定潜在对齐,提出了规范不可区分性概念,并通过实验验证了在评估意识下行为基准的局限性。

Comments 10 pages. Theoretical analysis of behavioral alignment evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06757 2026-01-13 cs.CL cs.AI 82%

MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues

MTMCS-Bench: 多轮对话中多模态大语言模型上下文安全性的评估

Zheyuan Liu, Dongwhi Kim, Yixin Wan, Xiangchi Yuan, Zhaoxuan Tan, Fengran Mo, Meng Jiang

机构 * University of Notre Dame(诺丁汉大学) University of California, Los Angeles(加州大学洛杉矶分校) Georgia Institute of Technology(佐治亚理工学院) University of Montreal(蒙特利尔大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

AI总结 MTMCS-Bench评估多模态大语言模型在多轮对话中的上下文安全性,揭示了安全与效用之间的权衡及现有防护措施的不足。

Comments A benchmark of realistic images and multi-turn conversations that evaluates contextual safety in MLLMs under two complementary settings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07170 2025-11-04 cs.LG cs.AI 82%

Trustworthy AI Must Account for Interactions

Jesse C. Cresswell

机构 * Jesse C. Cresswell(独立研究者)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG;alignment(comments)

Comments Presented at the ICLR 2025 Workshop on Bidirectional Human-AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07965 2025-04-11 cs.LG cs.CL 82%

Cat, Rat, Meow: On the Alignment of Language Model and Human Term-Similarity Judgments

Lorenz Linhardt, Tom Neuhäuser, Lenka Tětková, Oliver Eberle

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments ICLR 2025 Workshop on Representational Alignment (Re-Align)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01523 2025-03-03 cs.CL cs.AI 82%

GOAT-Bench: Safety Insights to Large Multimodal Models through Meme-Based Social Abuse

Hongzhan Lin, Ziyang Luo, Bo Wang, Ruichao Yang, Jing Ma

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments The first work to benchmark Large Multimodal Models in safety insight on social media

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15457 2025-02-13 cs.CY cs.AI cs.HC 82%

The Journey to Trustworthy AI: Pursuit of Pragmatic Frameworks

Mohamad M Nasr-Azadani, Jean-Luc Chatelain

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments Updates: Added disclaimer about USA's recent U-turn on Trustworthy AI Executive Order. Improved Fairness and Group size in section 6.5. Fixed typos. Added a few new references. Updated title

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05399 2025-01-13 cs.CL cs.AI 82%

SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety

Paul Röttger, Fabio Pernisi, Bertie Vidgen, Dirk Hovy

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI;alignment(comments)

Comments Accepted at AAAI 2025 (Special Track on AI Alignment)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01774 2024-01-08 cs.CY cs.AI cs.SE 82%

RE-centric Recommendations for the Development of Trustworthy(er) Autonomous Systems

Krishna Ronanki, Beatriz Cabrero-Daniel, Jennifer Horkoff, Christian Berger

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments Accepted at [TAS '23]{First International Symposium on Trustworthy Autonomous Systems}

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.07787 2022-02-17 cs.LG cs.AI 82%

Trustworthy Anomaly Detection: A Survey

Shuhan Yuan, Xintao Wu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments Paper list, see https://github.com/yuan-shuhan/trustworthy-anomaly-detection-papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28825 2026-05-29 cs.CL 81%

MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models

MechELK:一种用于激发大型语言模型中潜在知识的机制可解释性框架

Ji-jun Park, Soo-joon Choi, Jiwon Jeong, Taeyang Yoon, Ju-Wan Lee

机构 * Dongguk University(东国大学)

专题命中 安全评测 :alignment(abstract,abstract_cn);safety(abstract);AI safety(abstract);分类 cs.CL

AI总结 提出MechELK框架,通过定位、验证和激发三个阶段,利用稀疏自编码器特征分析和因果探测等方法,从大型语言模型中提取隐藏知识,在TruthfulQA等基准上平均激发准确率达84.7%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24826 2026-04-29 cs.CR cs.AI 81%

A Comparative Evaluation of AI Agent Security Guardrails

AI代理安全防护的比较评估

Qi Li, Jiu Li, Pingtao Wei, Jianjun Xu, Xueyi Wei, Jiwei Shi, Xuan Zhang, Yanhui Yang, Xiaodong Hui, Peng Xu, Lingquan Zhou

机构 * Beijing Caizhi Tech(北京彩智科技)

专题命中 安全评测 :safety(summary_cn,abstract);分类 cs.AI

AI总结 本文比较了DKnownAI Guard与AWS Bedrock Guardrails、Azure Content Safety和Lakera Guard在AI代理安全场景中的性能,发现DKnownAI在召回率和真正负样本率上表现最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04837 2026-03-06 cs.AI 81%

Design Behaviour Codes (DBCs): A Taxonomy-Driven Layered Governance Benchmark for Large Language Models

设计行为代码(DBCs):一种以分类为导向的分层治理基准用于大型语言模型

G. Madan Mohan, Veena Kiran Nambiar, Kiranmayee Janardhan

专题命中 安全评测 :alignment(abstract);RLHF(abstract);DPO(abstract);safety(abstract)

AI总结 本文提出DBC基准,用于评估大型语言模型中的行为治理层,通过三臂实验展示DBC在降低风险暴露率和提升合规性方面的效果。

Comments 14 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15859 2026-02-19 cs.CL 81%

From Transcripts to AI Agents: Knowledge Extraction, RAG Integration, and Robust Evaluation of Conversational AI Assistants

从语音记录到AI代理:知识提取、RAG整合及对话式AI助手的鲁棒评估

Krittin Pachtrachai, Petmongkon Pornpichitsuwan, Wachiravit Modecrua, Touchapon Kraisingkorn

机构 * Amity Research and Application Center (ARAC)(阿米蒂研究与应用中心)

专题命中 安全评测 :alignment(abstract);safety(abstract);red teaming(abstract);prompt injection(abstract)

AI总结 本文提出了一种端到端框架,通过历史通话记录构建和评估对话式AI助手,利用知识提取和RAG整合提升事实准确性和鲁棒性。

Comments 9 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08136 2026-02-10 cs.CV cs.AI 81%

Robustness of Vision Language Models Against Split-Image Harmful Input Attacks

视觉语言模型对分裂图像有害输入攻击的鲁棒性

Md Rafi Ur Rashid, MD Sadik Hossain Shanto, Vishnu Asutosh Dasu, Shagufta Mehnaz

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Bangladesh University of Engineering and Technology(孟加拉工程与技术大学)

专题命中 安全评测 :alignment(abstract);RLHF(abstract);safety(abstract);jailbreak(abstract)

AI总结 本研究提出分裂图像视觉陷阱攻击(SIVA),揭示视觉语言模型在面对分裂图像攻击时的安全漏洞,并通过对抗性知识蒸馏算法提升跨模型攻击效果。

Comments 22 Pages, long conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08494 2025-09-11 cs.CY cs.AI cs.CL cs.HC cs.LG 81%

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants

Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes, Jacy Reese Anthis

机构 * Apart Research AI Safety Cape Town University of Chicago(芝加哥大学) Stanford University(斯坦福大学) Sentience Institute(意识研究所)

专题命中 安全评测 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03053 2025-07-11 cs.MA cs.AI cs.CL cs.CY cs.LG 81%

MAEBE: Multi-Agent Emergent Behavior Framework

Sinem Erisken, Timothy Gothard, Martin Leitgab, Ram Potham

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Preprint. This work has been submitted to the Multi-Agent Systems Workshop at ICML 2025 for review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00911 2025-06-03 cs.AI 81%

Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language Models

William Overman, Mohsen Bayati

机构 * Graduate School of Business(商学院) Stanford University(斯坦福大学)

专题命中 安全评测 :alignment(abstract);safety(abstract);harmlessness(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14633 2025-05-21 cs.CL cs.AI cs.CY cs.HC cs.LG 81%

Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas

Yu Ying Chiu, Zhilin Wang, Sharan Maiya, Yejin Choi, Kyle Fish, Sydney Levine, Evan Hubinger

机构 * University of Washington(华盛顿大学) NVIDIA(NVIDIA公司) Cambridge(剑桥) Stanford(斯坦福大学) MIT(麻省理工学院) Harvard(哈佛大学) Anthropic(Anthropic公司)

专题命中 安全评测 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 34 pages, 11 figures, see associated data at https://huggingface.co/datasets/kellycyy/AIRiskDilemmas and code at https://github.com/kellycyy/LitmusValues

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13195 2025-05-20 cs.AI 81%

Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities

Lili Zhang, Haomiaomiao Wang, Long Cheng, Libao Deng, Tomas Ward

机构 * School of Computing, Dublin City University(都柏林城市大学计算机学院) Insight SFI Research Centre for Data Analytics(数据分析SFI研究中心) North China Electric Power University(华北电力大学) Harbin Institute of Technology (Weihai)(威海工业大学)

专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract);AI safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23497 2026-08-25 cs.AI cs.CL 新提交 81%

Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty

通过安全方向惩罚缓解推理诱导的不一致性

Yipeng Zhao, Qishun Yang, Shenzhe Zhu, Shu Yang, Di Wang

机构 * University of Toronto(多伦多大学) King Abdullah University of Science and Technology(阿卜杜拉国王科技大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

AI总结 针对推理微调引发LLM安全退化的RIM问题,本文提出安全方向惩罚(SDP)方法,通过定位安全决策层并惩罚安全方向位移,在Qwen2.5系列模型上实现安全与推理性能的平衡。

Comments 28 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22802 2026-08-25 cs.CL cs.AI 新提交 81%

SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support

面向可信临床决策支持的、关注社会决定健康因素(SDoH)的医学大语言模型叙事锚定偏差

Ahnaf Atef Choudhury, Ramkrishna Saha

机构 * George Mason University(乔治梅森大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

AI总结 该研究针对医学大语言模型的叙事锚定偏差问题,以Qwen2.5系列模型为对象开展实验,发现7B模型仍存在较高叙事敏感性误差,提出需结合正确率与叙事稳定性评估临床决策支持模型。

Comments Accepted for publication at 10th International Artificial Intelligence and Data Processing Symposium (IDAP'26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22335 2026-08-25 cs.CL cs.AI 新提交 81%

Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms

寄存器偏移破坏大语言模型安全:具有文化相关危害的孟加拉语基准测试

Naymul Islam, Nusrat Jahan Lia, Shubhashis Roy Dipta, Sabik Bin Sultan, Abdullah Khan Zehady

机构 * Institute of Information Technology, University of Dhaka(达卡大学信息技术学院) University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校) Bangladesh Air Force Shaheen College Kurmitola(孟加拉国空军库尔米托拉沙欣学院) Ciroos Inc.(Ciroos公司)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

AI总结 针对孟加拉语LLM安全评估的英语中心偏差问题,构建含879个提示的BanglaSafe基准,发现写作风格对有害请求成功率的影响最大,且现有安全分类器难以评估孟加拉语内容。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18164 2026-08-20 cs.CL cs.AI cs.CR 新提交 81%

Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation

M P V S Gopinadh

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments 3 pages. Accepted at ACL 2026 Workshop on Evaluation in Practice: Methodological Rigor, Sociotechnical Perspectives, & Community Collaboration (EvalEval)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16975 2026-08-19 cs.CL cs.LG 新提交 81%

Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence

用于脑-语言对应关系的边际正则化结构化语义对齐

Jiaqi Wang, Huawen Hu, Shu Zhang

机构 * School of Computer Science and Technology, Northwestern Polytechnical University(西北工业大学计算机科学与技术学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG

AI总结 本文针对脑-语言解码的模糊性问题,提出MD-SigLIP框架,通过边际正则化结构化语义对齐实现基于检索的解码,在全词汇与子集评估下均取得最优检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03361 2026-08-18 cs.CY cs.AI 版本更新 81%

The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk

价值的进化起源:对AI对齐、感知能力和生存风险的启示

Francis Heylighen

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.CY

AI总结 该研究通过追溯生物价值的进化起源,论证LLMs无内在生存动机与感知能力,正交性论点不适用于它们,AI对齐的核心挑战是确保LLMs应用习得的伦理价值。

Comments submitted chapter for book: T. Veloz & C. Rittberg (Eds.), AI and Human Values. Springer

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05177 2026-08-18 cs.CL cs.AI eess.AS 版本更新 81%

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

MCBench:面向全能大语言模型的多上下文安全评估基准

Manh Luong, Tamas Abraham, Junae Kim, Amar Kaur, Rollin Omari, Gholamreza Haffari, Trang Vu, Lizhen Qu, Dinh Phung

机构 * Monash University(墨尔本大学) Defence Science and Technology Group(国防科学与技术集团)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

AI总结 针对现有多模态安全基准仅处理视觉输入的局限,提出MCBench基准,包含1196个跨四类安全场景的测试,要求整合多模态信息进行安全评估,揭示当前全能大语言模型在跨模态安全推理上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06039 2026-08-18 cs.CL cs.AI 版本更新 81%

A Large-Scale Chinese Knowledge Graph-Text Alignment Dataset for Benchmarking Knowledge-Grounded LLMs

用于评估知识增强型大语言模型的大规模中文知识图谱-文本对齐数据集

Chengwei Wu, Xingrui Zhuo, Mingyang Gao, Xinghe Cheng, Zhichao Yan, Jiapu Wang

机构 * Nanjing University of Science and Techonolgy(南京理工大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) University of Science and Technology Beijing(北京科技大学) Hefei University of Technology(合肥工业大学) Beijing University of Chemical Technology(北京化工大学) Renmin University of China(中国人民大学) Beihang University(北航) Beijing University of Posts and Telecommunications(北京邮电大学) Griffith University, Australia(澳大利亚格里菲斯大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 该研究提出大规模中文知识图谱-文本对齐数据集CDTP,支持三项中文知识密集型任务的评估,经实验验证其可提升LLM的领域性能与分布外鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12360 2026-08-14 cs.CY cs.LG 新提交 81%

Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices

监管批准并不足够:FDA clearance医疗设备中可信赖AI报告的缺口

Ahmed M Salih, Oliver Díaz, Alejandro Guzman, Noah Marquez Vara, Fotios Avgoustidis, Rituraj Singh, Saman Barakat, Zahra Raisi-Estabragh, Karim Lekadir

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

AI总结 该研究分析2021-2025年FDA获批的519份AI/ML医疗设备报告,发现可信赖AI报告存在显著缺口,获批年份或临床领域不影响报告透明度,仅监管批准不足以证明AI可信赖性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12346 2026-08-14 cs.AI cs.CY 新提交 81%

Position: The Alignment Community is Unintentionally Building a Censor's Toolkit

立场:对齐社区正在无意之中构建审查者的工具包

Sarah Ball, Phil Hackemann

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.CY

AI总结 该立场论文指出现代AI对齐方法是两用技术,可能被恶意滥用,呼吁社区讨论其滥用风险并提出缓解策略。

Comments Accepted as oral paper at ICML 2026

Journal ref Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏