arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9400 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9400 篇

2008.01263 2020-08-05 cs.SE cs.AI cs.LG cs.SY eess.SY 81%

Safety design concepts for statistical machine learning components toward accordance with functional safety standards

Akihisa Morikawa, Yutaka Matsubara

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.03964 2020-07-09 math.OC cs.AI cs.LG 81%

Responsive Safety in Reinforcement Learning by PID Lagrangian Methods

Adam Stooke, Joshua Achiam, Pieter Abbeel

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments ICML 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20691 2026-07-24 cs.CV cs.AI 新提交 80%

Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis

用于可靠乳腺超声诊断的空间基础概念瓶颈模型

Moshiur Rahman Tonmoy, Dunren Che, Haitham Y. Adarbah, Afzel Noore

专题命中 安全评测 :trustworthy(title,comments);alignment(abstract);分类 cs.AI

AI总结 研究乳腺超声诊断中概念瓶颈模型可信度受监督限制问题,提出空间基础概念瓶颈模型(SG-CBM),利用病变轮廓弱监督,通过导出特定区域训练概念图,经交叉验证等提升诊断指标与概念证据空间对齐,强调数据质量监督设计及可信度验证的必要。

Comments Accepted to the Workshop on Data Quality Aware, High-Performance, and Trustworthy AI Systems for Healthcare at IEEE/ACM CHASE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25256 2026-06-05 cs.AI 80%

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

谁的对齐?比较不同组织决策情境下的大语言模型过程对齐

Niklas Weller, Emilio Barkett

机构 * University of Cambridge(剑桥大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出一种决策策略捕获方法测量过程对齐,发现LLM在ECHR第6条决策中过程对齐与输出准确性高度相关,但在德国消费信贷决策中关系消失,揭示了多元对齐挑战。

Comments Accepted to Pluralistic Alignment Workshop @ ICML 2026, Seoul, South Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19869 2026-05-20 cs.CV cs.AI 80%

Passive Construction Site Safety Monitoring via Persona-Scaffolded Adversarial Chain-of-Thought VLM Verification

通过基于人设的对抗性链式思考视觉语言模型验证实现被动施工现场安全监控

Ananth Sriram, Neel Mokaria, Rajveer Singh

机构 * Department of Computer Science, University of Maryland, College Park, MD, USA(大学马里兰学院计算机科学系,马里兰州科利尔帕克,MD,美国)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 本文提出了一种被动的施工现场安全监控方法,通过三阶段架构处理视频数据,结合细调的YOLO11、SAM 3和Qwen3-VL-8B-Instruct模型,利用基于人设的对抗性链式思考协议提高合规性验证和幻觉控制,主要贡献是第三阶段提示设计,提升了12%的精度。

Comments 10 pages, 4 figures. First place, Ironsite.ai Spatial Intelligence Hackathon, University of Maryland, February 2026. Code available at https://github.com/ananthsriram1/ironsite-hackathon-project-safety_assistant

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28087 2026-05-12 cs.LO cs.AI 80%

Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles

迈向基于法律和安全原则的神经符号因果规则合成、验证与评估

Zainab Rehan, Christian Medeiros Adriano, Sona Ghahremani, Holger Giese

机构 * Hasso Plattner Institute \ of Potsdam Prof.-Dr.-Helmert Str. 2-3, D-14482 Potsdam, Germany Hasso Plattner Institute \ of Potsdam

专题命中 安全评测 :safety(title,abstract);分类 cs.AI;trustworthy(journal_ref)

AI总结 本文提出一种神经符号因果框架,结合一阶逻辑抽象树、结构因果模型和深度强化学习,通过Meta层缓解目标误指定问题,实现可扩展的规则维护。

Journal ref Neurosymbolic eXplainable Trustworthy Systems @ AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19457 2026-04-22 cs.AI 80%

Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents

四轴决策对齐用于长周期企业AI代理

Vasundra Srininvasan

机构 * Vasundra Srinivasan

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出四轴对齐框架,用于评估长周期企业AI代理的决策行为,涵盖事实精度、推理连贯性、合规重建和校准回避,通过实验揭示了决策对齐的重要性。

Comments 21 pages, 5 figures, 8 tables. PDFLaTeX. Code and artifacts: https://github.com/vasundras/decision-alignment-long-horizon-agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18850 2026-04-22 cs.HC cs.AI cs.SI 80%

The Triadic Loop: A Framework for Negotiating Alignment in AI Co-hosted Livestreaming

三元循环:一种协商AI共播直播中对齐的框架

Katherine Wang, Nadia Berthouze, Aneesha Singh

机构 * University College London(伦敦大学学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出三元循环框架,用于协商AI共播直播中的对齐问题,通过三者之间的双向适应过程,解决多用户社交环境中的动态反馈循环问题。

Comments 6 pages, 1 figure, Proceedings the Human-AI Interaction Alignment Workshop at CHI 2026 (CHI26 BiAlign Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01816 2026-01-06 cs.AI 80%

Admissibility Alignment

可接受性对齐

Chris Duffey

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出MAP-AI架构,通过蒙特卡洛方法评估决策策略的可接受性,实现AI对齐的动态决策理论属性。

Comments 24 pages, 2 figures, 2 tables.. Decision-theoretic alignment under uncertainty

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16402 2025-11-21 cs.AI cs.DB 80%

Trustworthy AI in the Agentic Lakehouse: from Concurrency to Governance

可信AI在代理湖仓中的实现:从并发到治理

Jacopo Tagliabue, Federico Bianchi, Ciro Greco

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

AI总结 本文提出Bauplan设计,通过事务机制实现湖仓中的数据和计算隔离,解决代理工作流的可信性问题,并提供自修复管道的实现。

Comments AAAI26, pre-print of paper accepted at the Trustworthy Agentic AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15415 2025-06-19 cs.CL 80%

Targeted Lexical Injection: Unlocking Latent Cross-Lingual Alignment in Lugha-Llama via Early-Layer LoRA Fine-Tuning

Stanley Ngugi

机构 * Stanley Ngugi(独立研究者)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments 11 pages, 3 figures, 2 tables. Research on parameter-efficient fine-tuning (PEFT) for low-resource languages (Swahili). Investigates cross-lingual lexical alignment in Lugha-Llama using LoRA and contrastive learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09973 2024-11-18 cs.LG cs.IR 80%

Establishing and Evaluating Trustworthy AI: Overview and Research Challenges

Dominik Kowald, Sebastian Scher, Viktoria Pammer-Schindler, Peter Müllner, Kerstin Waxnegger, Lea Demelius, Angela Fessl, Maximilian Toller, Inti Gabriel Mendoza Estrada, Ilija Simic, Vedran Sabol, Andreas Truegler, Eduardo Veas, Roman Kern, Tomislav Nad, Simone Kopeinik

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments Accepted in Frontiers in Big Data and AI, Research Topic: Towards Fair AI for Trustworthy Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.06377 2023-08-14 cs.AI 80%

Relational Action Bases: Formalization, Effective Safety Verification, and Invariants (Extended Version)

Silvio Ghilardi, Alessandro Gianola, Marco Montali, Andrey Rivkin

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments Extended version of the conference paper 'Safety Verification and Universal Invariants for Relational Action Bases' by the same authors, accepted at the 32nd International Joint Conference on Artificial Intelligence (IJCAI 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10436 2023-04-21 cs.CL 80%

Safety Assessment of Chinese Large Language Models

Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, Minlie Huang

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments Benchmark website: http://coai.cs.tsinghua.edu.cn/leaderboard/ ; SafetyPrompts repo: https://github.com/thu-coai/Safety-Prompts

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.05657 2020-02-14 cs.CY 80%

Trustworthy AI in the Age of Pervasive Computing and Big Data

Abhishek Kumar, Tristan Braud, Sasu Tarkoma, Pan Hui

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY

Comments To be published in Percrowd 2020 (PerCom Adjunct). Please cite as: Abhishek Kumar, Tristan Braud, Sasu Tarkoma, Pan Hui. Trustworthy AI in the Age of Pervasive Computing and Big Data. In Proceedings of the 3rd International Workshop on Context-awareness for Multi-device Pervasive and Mobile Computing (Percrowd), Austin USA, March 2020 (Percom 2020 Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25982 2026-04-30 cs.LG cs.AI cs.CY cs.ET 80%

Open Problems in Frontier AI Risk Management

前沿人工智能风险管理中的开放问题

Marta Ziosi, Miro Plueckebaum, Stephen Casper, Henry Papadatos, Ze Shen Chin, Peter Slattery, James Gealy, Tim G. J. Rudner, Brian Tse, Ariel Gil, Patricia Paskov, Maximilian Negele, Rokas Gipiškis, Nada Madkour, Vera Lummis, Rupal Jain, Luise Eder, Kristina Fort, Malou C. van Draanen Glismann, Inès Belhadj, Amin Oueslati, Anna K. Wisakanto, Richard Mallah, Koen Holtman, Ranj Zuhdi, Daniel S. Schiff, Jessica Newman, Malcolm Murray, Robert Trager

机构 * Oxford Martin AI Governance Initiative, University of Oxford(牛津大学人工智能治理倡议) MIT Computer Science and Artificial Intelligence Laboratory, MIT(麻省理工学院计算机科学与人工智能实验室) MIT Future Tech(麻省理工学院未来技术) Stanford University(斯坦福大学) Governance and Responsible AI Lab, Purdue University(普渡大学治理与负责任的人工智能实验室) University of Toronto(多伦多大学) Mercatus Center, George Mason University(乔治·马歇尔大学麦卡锡中心) Vilnius University(维尔纽斯大学) Vijil SaferAI AI Standards Lab(人工智能标准实验室) The Future Society(未来社会) Concordia AI(康科德人工智能) Pivotal Research Center for AI Risk Management & Alignment(人工智能风险管理和对齐中心) UC Berkeley Center for Long-Term Cybersecurity(伯克利大学长期网络安全中心) Independent(独立)

专题命中 安全评测 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文探讨前沿人工智能风险管理中的核心问题,通过文献综述识别未解决的挑战,并分类问题类型以指导未来研究与治理。

Comments 81 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26354 2026-03-10 cs.AI cs.CL cs.LG 80%

Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents

你的代理可能误进化:自我进化大语言模型代理中的新兴风险

Shuai Shao, Qihan Ren, Chen Qian, Boyi Wei, Dadi Guo, Jingyi Yang, Xinhao Song, Linfeng Zhang, Weinan Zhang, Dongrui Liu, Jing Shao

专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究了自我进化代理中因自我进化偏离导致的误进化风险,揭示了其广泛存在及对安全对齐的影响,并提出缓解策略。

Comments Published in ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16497 2025-01-29 cs.LG cs.AI cs.CL cs.CR stat.ML 80%

Smoothed Embeddings for Robust Language Models

Ryo Hase, Md Rafi Ur Rashid, Ashley Lewis, Jing Liu, Toshiaki Koike-Akino, Kieran Parsons, Ye Wang

专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Presented in the Safe Generative AI Workshop at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19808 2025-01-22 cs.CL cs.AI cs.LG 80%

Can Models Learn Skill Composition from Examples?

Haoyu Zhao, Simran Kaur, Dingli Yu, Anirudh Goyal, Sanjeev Arora

专题命中 安全评测 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00399 2024-12-30 cs.CL cs.AI cs.LG 80%

Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code

Taishi Nakamura, Mayank Mishra, Simone Tedeschi, Yekun Chai, Jason T Stillerman, Felix Friedrich, Prateek Yadav, Tanmay Laud, Vu Minh Chien, Terry Yue Zhuo, Diganta Misra, Ben Bogin, Xuan-Son Vu, Marzena Karpinska, Arnav Varma Dantuluri, Wojciech Kusa, Tommaso Furlanello, Rio Yokota, Niklas Muennighoff, Suhas Pai, Tosin Adewumi, Veronika Laippala, Xiaozhe Yao, Adalberto Junior, Alpay Ariyak, Aleksandr Drozd, Jordan Clive, Kshitij Gupta, Liangyu Chen, Qi Sun, Ken Tsui, Noah Persaud, Nour Fahmy, Tianlong Chen, Mohit Bansal, Nicolo Monti, Tai Dang, Ziyang Luo, Tien-Tung Bui, Roberto Navigli, Virendra Mehta, Matthew Blumberg, Victor May, Huu Nguyen, Sampo Pyysalo

专题命中 安全评测 :safety(abstract);trustworthy(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13124 2025-01-24 cs.CL cs.AI 80%

Debate Helps Weak-to-Strong Generalization

Hao Lang, Fei Huang, Yongbin Li

专题命中 安全评测 :alignment(abstract,comments);safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments AAAI2025 Special Track on AI Alignment (Oral presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24621 2026-08-26 cs.CL 新提交 79%

Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding

超越语义准确率:面向安全关键型语言理解的后果感知评估

Yujing Chang, Thinh Pham, Van-Phat Thai, Chunyao Ma, Yash Guleria, Pham Nhut Huy, Sameer Alam

机构 * ATMRI, Nanyang Technological University (NTU)(南洋理工大学ATMRI) Centre of AI Research, VinUniversity(文大学人工智能研究中心) School of Management, Indian Institute of Technology Mandi(印度理工大学曼迪分校管理学院)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

AI总结 针对空管场景,研究发现传统NLP指标会误判语言模型的操作可靠性,提出后果感知评估可弥补语义-安全差距,为安全关键型部署提供必要补充。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24335 2026-08-26 cs.CL 新提交 79%

SteerCheck: Attribution Specificity and Alignment Leakage in Activation-Steering Audits

SteerCheck:激活引导审计中的归因特异性与对齐泄漏

Daming Luo, Christy Liang, Junyu Xuan

机构 * University of Technology Sydney(悉尼科技大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

AI总结 SteerCheck是一项预先注册的归因审计方法,用于检测激活引导的对齐泄漏,揭示了Qwen3-14B等模型在激活引导中的特异性局限,为审计相关结论提供了可操作的框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24252 2026-08-26 cs.AI cs.SE 新提交 79%

SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction

SA-Bench:评估基于大语言模型的论文复现中的语义对齐

Xue Hu, Zewei Pan, Zeli Su, Zhou Liu, Wentao Zhang

机构 * Beihang University(北京航空航天大学) Shanghai Jiao Tong University(上海交通大学) Minzu University of China(中央民族大学) Peking University(北京大学) Zhongguancun Academy(中关村科学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本研究推出SA-Bench基准,评估LLM复现论文时的语义对齐,发现现有模型复现准确率低,需优化脚手架以提升语义规范验证能力。

Comments Accepted to Findings of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23313 2026-08-25 cs.AI 新提交 79%

EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models

EviSafe:基于证据的视觉语言模型安全性评估

Xuetong Li, Gaofeng Liu

机构 * Department of Automation, Shanghai Jiao Tong University(上海交通大学自动化系)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 本研究提出基于证据的VLM安全性评估框架EviSafe及对应基准EviSafeBench,通过三探针协议评估11个VLMs,发现其未因正确多模态原因可靠安全,需开展超越拒绝次数的评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21898 2026-08-25 cs.AI 新提交 79%

Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning

训练需要可信的世界:用于智能体学习的经过验证的合成Web环境

Chenghao Zhang, Canran Xiao, SaiSai Hu, Dan Roth

机构 * University of Pennsylvania(宾夕法尼亚大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Pace University(佩斯大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

AI总结 该研究构建可执行、可审计的合成Web环境,修复缺陷后用于训练Web智能体,提升了可行任务率与PPO策略性能,增强了向多个基准的迁移能力,为智能体学习提供了可靠训练基底。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21880 2026-08-25 cs.CL cs.CR 新提交 79%

BanglaVeilGuard: Cross-Script Safety Benchmarking and Lightweight Guardrails for Bangla Large Language Models

BanglaVeilGuard:孟加拉语大语言模型的跨脚本安全基准测试与轻量安全防护措施

Md. Rakibul Hassan, Muhammad Iqbal Hossain

机构 * BRAC University(BRAC大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

AI总结 针对孟加拉语LLM跨脚本安全评估难题,本文提出BanglaVeilGuard基准与轻量提示防护,可降低攻击成功率,提升不安全请求召回率,但存在方言及带噪良性提示过度拒绝的问题。

Comments Accepted at the 4th International Conference on Computing Advancements (ICCA 2026). 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21057 2026-08-24 cs.LG 新提交 79%

Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

通过人类对齐设计用于药物发现中智能体AI的鲁棒性大语言模型评估系统

Emma Granqvist, Rocío Mercado, Samuel Genheden

机构 * AstraZeneca(阿斯利康) Chalmers University of Technology(查尔姆斯理工大学) University of Gothenburg(哥德堡大学) Science for Life Laboratory (SciLifeLab)(生命科学实验室)

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

AI总结 本研究针对药物发现智能体AI,提出人类对齐的LLM作为评判者评估框架,优化后对齐度达0.86,为科学领域智能体系统评估提供可复用模板。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14251 2026-08-24 cs.LG 版本更新 79%

Calibrate-Then-Delegate: Safety Monitoring with Risk and Budget Guarantees via Model Cascades

校准后再委托:通过模型级联实现具有风险和预算保证的安全监控

Edoardo Pona, Milad Kazemi, Mehran Hosseini, Yali Du, David Watson, Osvaldo Simeone, Nicola Paoletti

机构 * King’s College London(伦敦国王学院) University of Manchester(曼彻斯特大学) Northeastern University London(伦敦东北大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

AI总结 本文提出Calibrate-Then-Delegate方法,通过模型级联在保证计算成本的同时实现实例级决策,有效提升安全监控的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02022 2026-08-21 cs.AI 版本更新 79%

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

ATBench:一个多样且现实的代理轨迹基准,用于安全评估与诊断

Yu Li, Haoyu Luo, Yuejin Xie, Yuqian Fu, Zhonghao Yang, Shuai Shao, Qihan Ren, Wanying Qu, Yanwei Fu, Yujiu Yang, Jing Shao, Xia Hu, Dongrui Liu

机构 * Shanghai AI Lab(上海人工智能实验室) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) KAUST(卡塔尔人工智能科学中心) East China Normal University(华东师范大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 ATBench通过多样化的轨迹和现实场景评估代理安全,包含1000条轨迹,挑战性强,支持跨基准比较和长周期故障诊断。

详情

展开后加载摘要…

URL PDF HTML 收藏