arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9380 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9380 篇

1906.07838 2019-06-20 cs.LG cs.AI cs.RO stat.ML 62%

RadGrad: Active learning with loss gradients

Paul Budnarain, Renato Ferreira Pinto Junior, Ilan Kogan

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.02941 2019-05-09 cs.LG cs.AI stat.ML 62%

Robust Federated Training via Collaborative Machine Teaching using Trusted Instances

Yufei Han, Xiangliang Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.07640 2019-04-17 cs.CY cs.LG 62%

Medical device surveillance with electronic health records

Alison Callahan, Jason A Fries, Christopher Ré, James I Huddleston, Nicholas J Giori, Scott Delp, Nigam H Shah

专题命中 安全评测 :safety(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.02495 2019-04-05 cs.LG cs.AI 62%

A Categorisation of Post-hoc Explanations for Predictive Models

John Mitros, Brian Mac Namee

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 5 pages, 3 figures, AAAI 2019 Spring Symposia (#SSS19)

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.04442 2019-03-12 cs.AI cs.LG 62%

Physics Enhanced Artificial Intelligence

Patrick O'Driscoll, Jaehoon Lee, Bo Fu

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.04355 2018-08-14 cs.LG cs.AI cs.CV cs.RO stat.ML 62%

Large-Scale Study of Curiosity-Driven Learning

Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, Alexei A. Efros

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments First three authors contributed equally and ordered alphabetically. Website at https://pathak22.github.io/large-scale-curiosity/

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.10698 2018-06-29 cs.AI cs.LG stat.AP stat.ML 62%

A comparative study of artificial intelligence and human doctors for the purpose of triage and diagnosis

Salman Razzaki, Adam Baker, Yura Perov, Katherine Middleton, Janie Baxter, Daniel Mullarkey, Davinder Sangar, Michael Taliercio, Mobasher Butt, Azeem Majeed, Arnold DoRosario, Megan Mahoney, Saurabh Johri

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.08608 2017-03-06 stat.ML cs.AI cs.LG 62%

Towards A Rigorous Science of Interpretable Machine Learning

Finale Doshi-Velez, Been Kim

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.06997 2016-11-22 cs.CL cs.AI 62%

Coherent Dialogue with Attention-based Language Models

Hongyuan Mei, Mohit Bansal, Matthew R. Walter

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments To appear at AAAI 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1602.04938 2016-08-10 cs.LG cs.AI stat.ML 62%

"Why Should I Trust You?": Explaining the Predictions of Any Classifier

Marco Tulio Ribeiro, Sameer Singh, Carlos Guestrin

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.05547 2015-12-10 cs.CV cs.AI cs.LG cs.NE 62%

Return of Frustratingly Easy Domain Adaptation

Baochen Sun, Jiashi Feng, Kate Saenko

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Fixed typos. Full paper to appear in AAAI-16. Extended Abstract of the full paper to appear in TASK-CV 2015 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0006013 2009-11-30 cs.CL cs.AI 62%

An evaluation of Naive Bayesian anti-spam filtering

Ion Androutsopoulos, John Koutsias, Konstantinos V. Chandrinos, George Paliouras, Constantine D. Spyropoulos

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 9 pages

Journal ref Proceedings of the workshop on Machine Learning in the New Information Age, G. Potamias, V. Moustakis and M. van Someren (eds.), 11th European Conference on Machine Learning, Barcelona, Spain, pp. 9-17, 2000

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14329 2026-08-17 cs.CR cs.AI cs.CL cs.CY cs.LG 新提交 61%

A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation

面向基于原则的监管的LLM作为评判者的四轴可信度基准

Dipankar Sarkar

专题命中 安全评测 :分类 cs.CL、cs.AI、cs.CY;trustworthy(comments)

AI总结 本文提出面向基于原则监管的LLM评判者的四轴可信度基准,发布Principle-Bench并引入Ceca评估器,发现无方法在所有四轴占优,部署级LLM评判者需报告对抗性欺骗与事后校准等指标。

Comments 7 pages, 3 figures. Accepted at the KDD 2026 Workshop on Secure and Trustworthy Large Language Models (SeT-LLM), poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25423 2026-07-29 cs.HC cs.AI cs.ET 新提交 61%

From Dyad to Triad: Eliciting XAI Requirements in Stroke Rehabilitation

从二元组到三元组:中风康复中可解释人工智能需求的引出

Param Rajpura, Yogesh Kumar Meena

专题命中 安全评测 :trustworthy(abstract,comments);分类 cs.AI

AI总结 研究针对中风康复中引出可解释人工智能需求的挑战,提出基于视频的支架协议,包含类比桥接等四种方法,揭示了参与者的需求及引导偏差,为康复中引出面向患者的XAI需求提供可复用方法,是可靠人机系统设计前提。

Comments AI and Cognitive Computing for Trustworthy Human Machine Systems Session, IEEE International Conference on Systems, Man, and Cybernetics (IEEE SMC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14336 2026-07-17 cs.SE cs.AI 新提交 61%

Copy-on-Write Scoring: Application-Specific Agent Evaluations

写时复制评分:特定应用代理评估

Joanna Roy, Sven Hoelzel

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI;safety(comments)

AI总结 研究如何在软件系统中可靠部署基于大语言模型的代理,提出写时复制(CoW)评分框架,利用PostgreSQL级机制在应用环境中直接评估代理操作,能低成本评估迭代,在开源平台上验证了该框架的有效性。

Comments 15 pages, 11 figures, accepted at ICML 2026 Second Workshop on Agents in the Wild: Safety, Security, and Beyond

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29399 2026-06-30 cs.AI 61%

LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents

LLM引导的规划:多模态核监管文档的多跳推理

Mingyu Jeon, Bokyeong Kim, Suwan Cho, Jae Young Suh, Yonggyun Yu

专题命中 安全评测 :safety(abstract,comments);分类 cs.AI

AI总结 提出LLM引导的规划方法,通过动态知识图谱状态和文档树工具,在多跳推理任务中实现81.5%准确率,显著优于无状态规划方法。

Comments Accepted at the Second Workshop on Agents in the Wild: Safety, Security, and Beyond @ ICML 2026. 8 pages (main), 3 figures, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16167 2026-06-16 cs.AI 新提交 61%

AI Pluralism and the Worlds It Misses

AI多元主义及其遗漏的世界

Rashid Mushkani

机构 * Rashid Mushkani

专题命中 安全评测 :alignment(abstract,comments);分类 cs.AI

AI总结 本文提出AI系统施加本体论,导致本体论扁平化,并引入多元生命周期治理框架以记录本体开放性和问责条件。

Comments To be presented at the ICML Pluralistic Alignment Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15766 2026-06-16 cs.AI cs.HC 新提交 61%

Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments

重新思考LLM导师中的脚手架:基准测试与真实部署之间的交互不匹配

Alexandra Neagu, Jeffrey T. H. Wong, Marcus Messer, Rhodri Nelson, Peter B. Johnson

机构 * University of Cambridge(剑桥大学)

专题命中 安全评测 :alignment(abstract,comments);分类 cs.AI

AI总结 通过分析9490个聊天记录,发现AI导师基准测试假设学生积极接受脚手架,但真实场景中学生常绕过脚手架,揭示基准测试与真实部署的交互不匹配。

Comments Pluralistic Alignment Workshop @ ICML 2026, Seoul, South Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14347 2026-06-15 cs.LG 新提交 61%

When Language Representations Interact: Separability and Cross-Lingual Effects in LLMs

当语言表示交互时:LLM中的可分离性与跨语言效应

Boris Marinov, Angira Sharma, Christian Schroeder de Witt, Philip Torr, Anisoara Calinescu, Jialin Yu

机构 * University of Oxford(牛津大学) Imperial College London(帝国理工学院) University of York(约克大学)

专题命中 安全评测 :trustworthy(abstract,comments);分类 cs.LG

AI总结 通过因果几何分析,研究多语言LLM中语言表示的线性可分离性及跨语言结构依赖,发现语言概念在协方差调整内积下可分离,同语系语言呈现单纯形几何结构。

Comments Trustworthy AI for Good (AI4Good) Workshop @ ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16651 2026-06-11 cs.CV cs.LG 版本更新 61%

Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations

正确预测,误导性解释:关于视觉-语言模型解释的脆弱性

Narges Babadi, Hadis Karimipour

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(abstract);分类 cs.LG;trustworthy(comments)

AI总结 研究探讨了视觉-语言模型中解释热图在对抗条件下是否忠实反映推理过程,提出X-Shift攻击揭示解释与预测行为的脱节,验证了解释机制的脆弱性。

Comments Accepted at the ICML 2026 Workshop on Trustworthy AI for Good (AI4GOOD), Seoul, South Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24229 2026-05-26 cs.AI 61%

How Well Do Models Follow Their Constitutions?

模型遵循其宪法的程度如何?

Arya Jakkli, Senthooran Rajamanoharan, Neel Nanda

机构 * Anthropic

专题命中 安全评测 :alignment(abstract,comments);分类 cs.AI

AI总结 提出多方法审计流程,评估前沿AI模型在对抗性多轮交互中遵循其书面行为规范(如Anthropic宪法和OpenAI模型规范)的程度,发现新一代模型违规率显著下降。

Comments 37 pages including appendix. Code, tenet lists, and full transcripts: https://github.com/ajobi-uhc/constitution-audits. Companion blog post on LessWrong/AI Alignment Forum: https://www.lesswrong.com/posts/Tk4SF8qFdMrzGJGGw/how-well-do-models-follow-their-constitutions

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13880 2026-03-17 cs.NE cs.AR cs.LG cs.RO 61%

Benchmarking the Energy Cost of Assurance in Neuromorphic Edge Robotics

在神经形态边缘机器人中评估保证的能量成本

Sylvester Kaczmarek

专题命中 安全评测 :trustworthy(abstract,comments);分类 cs.LG

AI总结 本文研究了在边缘机器人中部署可信人工智能的高保证鲁棒性和能源可持续性之间的权衡,通过对比传统防御机制与事件驱动的神经形态架构,展示了后者在降低攻击成功率的同时保持低能耗的优势。

Comments 6 pages, 4 figures. Accepted and presented at the STEAR 2026 Workshop on Sustainable and Trustworthy Edge AI for Robotics, HiPEAC 2026, Krakow, Poland

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13325 2026-03-17 cs.MA cs.AI 61%

Auditing Cascading Risks in Multi-Agent Systems via Semantic-Geometric Co-evolution

通过语义-几何共演化审计多智能体系统中的级联风险

Zixun Luo, Yuhang Fan, Hengyu Lin, Yufei Li, Youzhi Zhang

机构 * Huazhong University of Science and Technology(华中科技大学) Lingnan University(岭大大学) Tsinghua University(清华大学) Centre for Artificial Intelligence and Robotics(人工智能与机器人中心) Hong Kong Institute of Science and Innovation(香港创新科技研究院) Chinese Academy of Sciences(中国科学院)

专题命中 安全评测 :trustworthy(abstract,comments);分类 cs.AI

AI总结 本文提出基于语义-几何共演化的框架,通过动态图模型和Ollivier-Ricci曲率检测多智能体系统中的级联风险,提前预警并定位根本原因。

Comments This work has been accepted to ICLR 2026 Workshop: Principled Design for Trustworthy AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05399 2026-03-06 cs.AI 61%

Judge Reliability Harness: Stress Testing the Reliability of LLM Judges

LLM裁判可靠性检测工具:检验LLM裁判的可靠性

Sunishchal Dev, Andrew Sloan, Joshua Kavner, Nicholas Kong, Morgan Sandler

机构 * RAND Corporation(RAND公司)

专题命中 安全评测 :safety(abstract,comments);分类 cs.AI

AI总结 本研究提出了一种开源工具,用于评估LLM裁判在不同基准中的可靠性,发现模型和扰动类型对性能有显著影响。

Comments Accepted at Agents in the Wild: Safety, Security, and Beyond Workshop at ICLR 2026 - April 26, 2026, Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05264 2026-01-12 cs.IR cs.AI 61%

Engineering the RAG Stack: A Comprehensive Review of the Architecture and Trust Frameworks for Retrieval-Augmented Generation Systems

构建RAG堆栈:对检索增强生成系统架构和信任框架的全面综述

Dean Wampler, Dave Nielson, Alireza Seddighi

机构 * The AI Alliance(AI联盟) IBM Research(IBM研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI;safety(comments)

AI总结 本文综述了RAG系统架构和信任框架,提出统一分类学和评估框架,为构建安全且领域适应性强的RAG系统提供指导。

Comments 86 pages, 2 figures, 37 tables. A comprehensive review of Retrieval-Augmented Generation (RAG) architectures and trust frameworks (2018-2025), encompassing a unified taxonomy, evaluation benchmarks, and trust-safety modeling

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11393 2025-11-24 cs.SE cs.AI cs.CR cs.MA 61%

LLM-Agent-UMF: LLM-based Agent Unified Modeling Framework for Seamless Design of Multi Active/Passive Core-Agent Architectures

LLM-Agent-UMF: 基于大语言模型的代理统一建模框架,用于无缝设计多主动/被动核心代理架构

Amine Ben Hassouna, Hana Chaari, Ines Belhaj

机构 * Mediterranean Institute of Technology(地中海技术研究所) South Mediterranean University(南地中海大学) National School of Computer Science(国家计算机科学学校) University of Manouba(曼努巴大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI;trustworthy(comments)

AI总结 LLM-Agent-UMF提出了一种基于大语言模型的代理统一建模框架,用于设计多主动/被动核心代理架构,解决了现有架构的模块化和术语不一致问题。

Comments 39 pages, 19 figures, 3 tables. Published in Information Fusion, Volume 127, March 2026, 103865. Part of the special issue "Data Fusion Approaches in Data-Centric AI for Developing Trustworthy AI Systems"

Journal ref Information Fusion 127 (2026) 103865

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02879 2025-04-09 cs.CL 61%

Position: LLM Unlearning Benchmarks are Weak Measures of Progress

Pratiksha Thaker, Shengyuan Hu, Neil Kale, Yash Maurya, Zhiwei Steven Wu, Virginia Smith

专题命中 安全评测 :safety(abstract);分类 cs.CL;trustworthy(comments)

Comments Appears in IEEE Secure and Trustworthy Machine Learning (SaTML) '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03731 2025-04-08 cs.AI 61%

A Benchmark for Scalable Oversight Protocols

Abhimanyu Pallavi Sudhir, Jackson Kaunismaa, Arjun Panickssery

专题命中 安全评测 :alignment(abstract,comments);分类 cs.AI

Comments Accepted at the ICLR 2025 Workshop on Bidirectional Human-AI Alignment (BiAlign)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14814 2025-03-21 eess.SY cs.CV cs.LG cs.SY 61%

Probabilistic Risk Assessment of an Obstacle Detection System for GoA 4 Freight Trains

Mario Gleirscher, Anne E. Haxthausen, Jan Peleska

专题命中 安全评测 :safety(abstract,journal_ref);分类 cs.LG

Journal ref 9th ACM SIGPLAN International Workshop on Formal Techniques for Safety-Critical Systems (FTSCS), 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01318 2024-11-04 cs.CR cs.LG 61%

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tramer, Hamed Hassani, Eric Wong

专题命中 安全评测 :jailbreak(abstract,comments);分类 cs.LG

Comments The camera-ready version of JailbreakBench v1.0 (accepted at NeurIPS 2024 Datasets and Benchmarks Track): more attack artifacts, more test-time defenses, a more accurate jailbreak judge (Llama-3-70B with a custom prompt), a larger dataset of human preferences for selecting a jailbreak judge (300 examples), an over-refusal evaluation dataset, a semantic refusal judge based on Llama-3-8B

详情

展开后加载摘要…

URL PDF HTML 收藏