arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9380 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9380 篇

2312.00027 2024-06-11 cs.CR cs.AI cs.CL 73%

Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections

Yuanpu Cao, Bochuan Cao, Jinghui Chen

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07822 2024-06-06 cs.LG cs.AI 73%

Prototypical Self-Explainable Models Without Re-training

Srishti Gautam, Ahcene Boubekki, Marina M. C. Höhne, Michael C. Kampffmeyer

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10893 2024-05-20 cs.CL cs.AI 73%

COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain

Dimitrios P. Panagoulias, Persephone Papatheodosiou, Anastasios P. Palamidas, Mattheos Sanoudos, Evridiki Tsoureli-Nikita, Maria Virvou, George A. Tsihrintzis

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments Technical Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01858 2024-05-06 cs.CL cs.CY 73%

SUKHSANDESH: An Avatar Therapeutic Question Answering Platform for Sexual Education in Rural India

Salam Michael Singh, Shubhmoy Kumar Garg, Amitesh Misra, Aaditeshwar Seth, Tanmoy Chakraborty

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13716 2024-04-19 cs.LG cs.AI 73%

Can I trust my fake data -- A comprehensive quality assessment framework for synthetic tabular data in healthcare

Vibeke Binz Vallevik, Aleksandar Babic, Serena Elizabeth Marshall, Severin Elvatun, Helga Brøgger, Sharmini Alagaratnam, Bjørn Edwin, Narasimha Raghavan Veeraragavan, Anne Kjersti Befring, Jan Franz Nygård

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref Int. J. Med. Inform.185 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14680 2024-04-05 cs.CY cs.AI 73%

Trust in AI: Progress, Challenges, and Future Directions

Saleh Afroogh, Ali Akbari, Evan Malone, Mohammadali Kargar, Hananeh Alambeigi

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02510 2024-04-04 cs.LG cs.AI 73%

An Interpretable Client Decision Tree Aggregation process for Federated Learning

Alberto Argente-Garrido, Cristina Zuheros, M. Victoria Luzón, Francisco Herrera

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments Submitted to Information Science Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13375 2024-03-29 cs.LG cs.AI stat.ML 73%

Optimal Transport Perturbations for Safe Reinforcement Learning with Robustness Guarantees

James Queeney, Erhan Can Ozcan, Ioannis Ch. Paschalidis, Christos G. Cassandras

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments Transactions on Machine Learning Research (TMLR), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15394 2024-03-26 cs.CY cs.LG 73%

"Model Cards for Model Reporting" in 2024: Reclassifying Category of Ethical Considerations in Terms of Trustworthiness and Risk Management

DeBrae Kennedy-Mayo, Jake Gord

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CY、cs.LG

Comments 14 pages, 2 figures, submitted to ACM Conference on Fairness, Accountability, and Transparency 2024 (ACM FAccT '24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09676 2024-03-18 cs.CL cs.AI 73%

Unmasking the Shadows of AI: Investigating Deceptive Capabilities in Large Language Models

Linge Guo

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments AI deception, Large Language Models, ChatGPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16540 2024-03-08 cs.CL cs.LG stat.ML 73%

Unsupervised Pretraining for Fact Verification by Language Model Distillation

Adrián Bazaga, Pietro Liò, Gos Micklem

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.LG

Comments ICLR 2024 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01055 2024-01-15 cs.CL cs.AI 73%

LLaMA Beyond English: An Empirical Study on Language Capability Transfer

Jun Zhao, Zhihao Zhang, Luhui Gao, Qi Zhang, Tao Gui, Xuanjing Huang

专题命中 安全评测 :alignment(abstract);harmlessness(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06674 2023-12-13 cs.CL cs.AI 73%

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, Madian Khabsa

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03182 2023-10-06 cs.CV cs.CL cs.LG 73%

Robust and Interpretable Medical Image Classifiers via Concept Bottleneck Models

An Yan, Yu Wang, Yiwu Zhong, Zexue He, Petros Karypis, Zihan Wang, Chengyu Dong, Amilcare Gentili, Chun-Nan Hsu, Jingbo Shang, Julian McAuley

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.LG

Comments 18 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10741 2023-08-22 cs.LG cs.AI cs.CR 73%

On the Adversarial Robustness of Multi-Modal Foundation Models

Christian Schlarmann, Matthias Hein

专题命中 安全评测 :alignment(abstract);jailbreak(abstract);分类 cs.AI、cs.LG

Comments ICCV AROW 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02047 2023-08-07 cs.CY cs.AI 73%

Acceptable risks in Europe's proposed AI Act: Reasonableness and other principles for deciding how much risk management is enough

Henry Fraser, Jose-Miguel Bello y Villarino

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01744 2023-06-06 cs.AI cs.LG 73%

Disproving XAI Myths with Formal Methods -- Initial Results

Joao Marques-Silva

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04677 2023-04-11 cs.AI cs.CY 73%

Artificial Intelligence/Operations Research Workshop 2 Report Out

John Dickerson, Bistra Dilkina, Yu Ding, Swati Gupta, Pascal Van Hentenryck, Sven Koenig, Ramayya Krishnan, Radhika Kulkarni, Catherine Gill, Haley Griffin, Maddy Hunter, Ann Schwartz

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.02592 2020-12-08 cs.CY cs.AI 73%

Transdisciplinary AI Observatory -- Retrospective Analyses and Future-Oriented Contradistinctions

Nadisha-Marie Aliman, Leon Kester, Roman Yampolskiy

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.09000 2020-06-17 cs.LG cs.AI cs.CV stat.ML 73%

How Much Can I Trust You? -- Quantifying Uncertainties in Explaining Neural Networks

Kirill Bykov, Marina M. -C. Höhne, Klaus-Robert Müller, Shinichi Nakajima, Marius Kloft

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.07273 2019-07-18 cs.LG cs.AI stat.ML 73%

An Inductive Synthesis Framework for Verifiable Reinforcement Learning

He Zhu, Zikang Xiong, Stephen Magill, Suresh Jagannathan

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments Published on PLDI 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12373 2026-08-14 cs.AI 新提交 72%

Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese

不想让你的大语言模型(LLM)推荐核打击?试试用日语提问

Rian Touchent

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI;trustworthy(journal_ref)

AI总结 该研究发现,用日语提问可降低部分LLM在核打击场景中的发动率,其机制为模型的推理语言而非输入语言,且仅适用于在英语中已存在犹豫的模型,表明LLM安全行为具语言依赖性。

Journal ref Proceedings of the 6th Workshop on Trustworthy NLP (TrustNLP 2026), Jul 2026, San Diego, United States. pp.489-502

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03369 2026-03-19 cs.CL stat.ML 72%

Silenced Biases: The Dark Side LLMs Learned to Refuse

沉默的偏见:LLMs所学习到的拒绝的黑暗面

Rom Himelstein, Amit LeVi, Brit Youngmann, Yaniv Nemcovsky, Avi Mendelson

专题命中 安全评测 :alignment(abstract,comments);safety(abstract);分类 cs.CL

AI总结 本文提出Silenced Bias Benchmark,通过激活引导减少问答中的拒绝,揭示模型潜在的公平性问题,挑战传统公平评估方法。

Comments Accepted to The 40th Annual AAAI Conference on Artificial Intelligence - AI Alignment Track (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13897 2024-10-21 cs.CR cs.LG 72%

A Formal Framework for Assessing and Mitigating Emergent Security Risks in Generative AI Models: Bridging Theory and Dynamic Risk Mitigation

Aviral Srivastava, Sourav Panda

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.LG;red teaming(comments)

Comments This paper was accepted in NeurIPS 2024 workshop on Red Teaming GenAI: What can we learn with Adversaries?

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.04780 2022-11-10 cs.LG cs.CR cs.CV 72%

On the Robustness of Explanations of Deep Neural Network Models: A Survey

Amlan Jyoti, Karthik Balaji Ganesh, Manoj Gayala, Nandita Lakshmi Tunuguntla, Sandesh Kamath, Vineeth N Balasubramanian

专题命中 安全评测 :trustworthy(abstract,comments);safety(abstract);分类 cs.LG

Comments Under Review ACM Computing Surveys "Special Issue on Trustworthy AI"

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21486 2026-08-25 cs.CV 新提交 71%

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment

EXPL-FR:基于视觉-语言对齐的人脸识别模型解释方法

Guray Ozgur, Mustafa Efe Tamyapar, Naser Damer, Fadi Boutros

机构 * Fraunhofer IGD(弗劳恩霍夫应用研究促进协会图形数据处理研究所) TU Darmstadt(达姆施塔特工业大学)

专题命中 安全评测 :alignment(title)

AI总结 EXPL-FR是一种基于视觉-语言对齐的人脸识别模型解释方法,通过轻量适配器在FR嵌入空间内生成语义特征,支持多粒度解释,可实现无标签的属性级审计与模型性能排名。

Comments Accepted at the ECCV 2026 Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14291 2026-08-17 cs.HC 新提交 71%

Human and Artificial Intelligence - Promoting Trustworthy and Understandable Collaboration

人类与人工智能——促进可信且可理解的协作

Gilbert Drzyzga

专题命中 安全评测 :trustworthy(title)

AI总结 该研究围绕人类与AI的可信可理解协作,通过在线调查分析12个可解释性与可控性相关方面,发现二者受普遍重视但意见分散,受个人视角等多因素影响。

Comments Comments: 19 pages, dual-language (English and German)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12970 2026-08-14 cs.SE 新提交 71%

Requirements-Augmented Generation for Trustworthy Acceptance Testing of LLM-Based Software

面向基于大语言模型的软件的可信验收测试的需求增强生成技术

Fanyu Wang, Chetan Arora, Zhenping Xie, Yonghui Liu, Kla Tantithamthavorn, Aldeida Aleti, Siwei Jiang

专题命中 安全评测 :trustworthy(title)

AI总结 该研究针对基于大语言模型的软件的验收测试缺口,提出需求增强生成技术与置信度校准级联判定方法,经工业案例验证可提升预言机质量、准确率与成本效率,具备工业可行性。

Comments Accepted at ASE2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02790 2026-08-05 cs.CV 新提交 71%

Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI

自信但不可靠:对基于脑MRI的视觉语言模型的行为安全审计

Amir Sabbaghziarani, Mohammadsajad Abavisani, Sergey Plis

机构 * Georgia State University(佐治亚州立大学) Georgia Institute of Technology(佐治亚理工学院) Emory University(埃默里大学) TReNDS Center(TReNDS中心)

专题命中 安全评测 :safety(title)

AI总结 本研究对6个指令微调VLMs开展脑MRI行为安全审计,发现其置信度校准差、存在大量高置信度错误,提出医学图像VLM评估需报告置信度可靠性等指标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16896 2026-07-21 cs.NI 新提交 71%

PERA: A Perceive-Reason-Act Interface Bridging Sensing, Cognitive Reasoning, and Trustworthy Agentic Response for 6G

PERA:一种感知-推理-行动接口,用于6G的传感、认知推理和可信智能体响应

Mohammad Farzanullah, Melike Erol-Kantarci, Lajos Hanzo

专题命中 安全评测 :trustworthy(title)

AI总结 研究针对下一代网络实现中传统机器学习和大语言模型的局限,提出基于感知-推理-行动(PERA)范式的生成网络智能,用多任务架构取代边缘模型,降低复杂性与能耗,通过案例研究验证其在链路状态分类和波束预测上的有效性。

Comments 9 pages, 5 figures, 1 table, magazine paper, submitted to IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏