arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8017 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8017 篇

1909.03464 2019-11-25 cs.CL stat.ML 79%

Back to the Future -- Sequential Alignment of Text Representations

Johannes Bjerva, Wouter Kouw, Isabelle Augenstein

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments AAAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.04706 2019-11-22 cs.SE cs.LG 79%

Towards Safety Verification of Direct Perception Neural Networks

Chih-Hong Cheng, Chung-Hao Huang, Thomas Brunner, Vahid Hashemi

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments Revised text (2nd version). The research work is conducted during the first author's service at the fortiss research institute and is supported by the following projects: "Audi Verifiable AI" from Audi AG, Germany and "Dependable AI for automotive systems" from DENSO Corporation, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.11768 2019-11-05 stat.ML cs.LG 79%

Hierarchical Optimal Transport for Multimodal Distribution Alignment

John Lee, Max Dabagia, Eva L. Dyer, Christopher J. Rozell

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.07658 2019-10-15 stat.ML cs.LG 79%

Graph-based regularization for regression problems with alignment and highly-correlated designs

Yuan Li, Benjamin Mark, Garvesh Raskutti, Rebecca Willett, Hyebin Song, David Neiman

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.06746 2019-07-30 cs.LG stat.ML 79%

nn-dependability-kit: Engineering Neural Networks for Safety-Critical Autonomous Driving Systems

Chih-Hong Cheng, Chung-Hao Huang, Georg Nührenberg

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments Tool available at https://github.com/dependable-ai/nn-dependability-kit

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.09347 2019-04-24 cs.LG cs.CV stat.ML 79%

Joint Domain Alignment and Discriminative Feature Learning for Unsupervised Deep Domain Adaptation

Chao Chen, Zhihong Chen, Boyuan Jiang, Xinyu Jin

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments This paper has been accepted by AAAI-2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.05464 2019-04-03 cs.LG cs.HC q-bio.NC stat.ML 79%

Transfer Learning for Brain-Computer Interfaces: A Euclidean Space Data Alignment Approach

He He, Dongrui Wu

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.04268 2018-04-13 cs.AI 79%

Incomplete Contracting and AI Alignment

Dylan Hadfield-Menell, Gillian Hadfield

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.03449 2017-01-13 stat.ML cs.LG math.PR 79%

Manifold Alignment Determination: finding correspondences across different data views

Andreas Damianou, Neil D. Lawrence, Carl Henrik Ek

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments NIPS workshop on Multi-Modal Machine Learning, 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.04576 2016-10-17 cs.LG stat.ML 79%

Kernel Alignment Inspired Linear Discriminant Analysis

Shuai Zheng, Chris Ding

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Joint European Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD, 2014

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.03650 2016-04-28 cs.CL 79%

Smoothing parameter estimation framework for IBM word alignment models

Vuong Van Bui, Cuong Anh Le

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1109.6838 2011-10-03 cs.MA cs.AI 79%

Distributed Air Traffic Control : A Human Safety Perspective

Sarvesh Nikumbh, Joeprakash Nathaman, Rahul Vartak

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Comments Extended Abstract, 3 pages, Accepted at IBM Collaborative Academia Research Exchange (I-CARE)-2011, uses ACM-Proceeding style file

详情

展开后加载摘要…

URL PDF HTML 收藏
1106.4570 2011-06-24 cs.GT cs.AI 79%

Competitive Safety Analysis: Robust Decision-Making in Multi-Agent Systems

M. Tennenholtz

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Journal ref Journal Of Artificial Intelligence Research, Volume 17, pages 363-378, 2002

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0311045 2009-12-01 cs.AI 79%

Unsupervised Grammar Induction in a Framework of Information Compression by Multiple Alignment, Unification and Search

J Gerard Wolff

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Journal ref Proceedings of the Workshop and Tutorial on Learning Context-Free Grammars (in association with the 14th European Conference on Machine Learning and the 7th European Conference on Principles and Practice of Knowledge Discovery in Databases (ECML/PKDD 2003), September 2003, Cavtat-Dubrovnik, Croata), editors: C. de la Higuera and P. Adriaans and M. van Zaanen and J. Oncina, pp 113-124

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09702 2026-07-14 cs.GT q-fin.TR 新提交 79%

Fundamental market design as a layer of AI-agent alignment

作为人工智能-智能体对齐层次的基础市场设计

Omar Inverso, Emilio Tuosto, Dragisa Zunic

专题命中 其他安全 :alignment(title,abstract)

AI总结 研究探讨市场中人工智能-智能体对齐,提出将基础市场设计视为该对齐层次,通过对市场核心形式化建模,利用理论计算机科学严谨性构建透明盒模型,支持激励分析与机制设计,使期望行为受青睐,不良行为难维持。

Comments Accepted as at the EC'26 Workshop on Incentive-Based AI Alignment, co-located with the 27th ACM Conference on Economics and Computation, Rome, Italy, July 2026. This version is prepared for public dissemination following workshop acceptance

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15300 2026-05-18 cs.CV 79%

Deep Pre-Alignment for VLMs

视觉语言模型的深度预对齐

Tianyu Yu, Kechen Fang, Zihao Wan, Kaidong Zhang, Yicheng Zhang, Jun Song, Bo Zheng, Yuan Yao

机构 * Tsinghua University Shanghai Qi Zhi Institute Taobao \& Tmall Group of Alibaba

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文提出深度预对齐(DPA),通过替换传统ViT编码器为小型VLM作为感知器,实现视觉特征与目标大语言模型文本空间的深度对齐,提升了多模态基准性能,并降低了语言能力遗忘。

Comments Accepted by ICML 2026. Project Website: https://github.com/THUMAI-Lab/Deep-Pre-Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04295 2024-12-17 cs.CV 79%

Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment

Jiayi Guo, Junhao Zhao, Chaoqun Du, Yulin Wang, Chunjiang Ge, Zanlin Ni, Shiji Song, Humphrey Shi, Gao Huang

专题命中 其他安全 :alignment(title,abstract)

Comments GitHub: https://github.com/SHI-Labs/Diffusion-Driven-Test-Time-Adaptation-via-Synthetic-Domain-Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13266 2024-10-18 physics.soc-ph 79%

Continuous agent-based modeling of adult-child pairs based on a pseudo-energy: Relevance for public safety and egress efficiency

Chuan-Zhi Thomas Xie, Tie-Qiao Tang, Alexandre Nicolas

专题命中 其他安全 :safety(title,abstract)

Journal ref Safety Science, 2024, 177, pp.106576

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07885 2022-12-21 cs.CV 79%

Clover: Towards A Unified Video-Language Alignment and Fusion Model

Jingjia Huang, Yinan Li, Jiashi Feng, Xinglong Wu, Xiaoshuai Sun, Rongrong Ji

专题命中 其他安全 :alignment(title,abstract)

Comments Update Tri-modal Alignment task

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22826 2026-05-25 cs.CL cs.AI cs.GT cs.MA 79%

Evaluating Large Language Models in a Complex Hidden Role Game

评估大型语言模型在复杂隐藏角色游戏中的表现

Niklas Bauer

机构 * University of Göttingen(哥廷根大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过社交推理游戏《秘密希特勒》评估大型语言模型的推理、说服和欺骗能力,引入新指标并发现当前模型在复杂多轮操纵中效果不佳。

Comments Master's thesis, University of Göttingen

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06696 2026-05-11 cs.AI cs.LG cs.MA 79%

Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations

多智能体AI中的隐藏联盟:来自内部表示的谱诊断

Cameron Berg, Susan L. Schneider, Mark M. Bailey

机构 * Reciprocal Research(递归研究) Center for the Future of AI, Mind, and Society(人工智能、心智与社会未来中心) Florida Atlantic University(佛罗里达 Atlantic 大学) Biological and Computational Intelligence Center(生物与计算智能中心) National Intelligence University(国家情报大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过分析多智能体系统内部神经表示的谱分区方法,检测隐藏联盟结构,验证了该方法在强化学习和大语言模型中的有效性,揭示了代表层次结构。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16448 2026-03-31 cs.AI cs.LG 79%

Information-theoretic Distinctions Between Deception and Confusion

信息论视角下欺骗与混淆的区别

Robin Young

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 本文从信息论角度区分两种AI安全失效模式:欺骗对齐与目标漂移,揭示二者在人类-AI系统不同接口的信息分歧,提出形式化模型和思想实验,为大型语言模型对齐挑战提供新视角。

Comments Proceedings of the 14th IJCNLP and the 4th AACL (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17577 2026-01-27 cs.HC cs.AI cs.CL 79%

Status Hierarchies in Language Models

语言模型中的地位层级

Emilio Barkett

机构 * Brigham Young University–Hawaii COLUMBIA UNIVERSITY(哥伦比亚大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

AI总结 语言模型在多智能体环境中会因地位线索形成层级,高地位分配反而降低高能力模型的服从,揭示AI系统中的新兴社会行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03442 2025-11-17 cs.CL cs.AI 79%

Are language models rational? The case of coherence norms and belief revision

Thomas Hofweber, Peter Hase, Elias Stengel-Eskin, Mohit Bansal

机构 * Department of Philosophy University of North Carolina at Chapel Hill(哲学系北卡罗来纳大学教堂山分校) Department of Computer Science University of North Carolina at Chapel Hill(计算机科学系北卡罗来纳大学教堂山分校) Department of Computer Science University of Texas at Austin(计算机科学系德克萨斯大学奥斯汀分校)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments substantial expansions of sections 4 and 5, updated references, numerous smaller additions and clarifications

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05580 2025-11-11 q-bio.NC cs.CL cs.CY cs.HC 79%

Approximating the Mathematical Structure of Psychodynamics

Bryce-Allen Bagley, Navin Khoshnan

机构 * Stanford University(斯坦福大学) Mathematical Medicine Group(数学医学组) Department of Neurosurgery(神经外科系) Physician-Scientist Training Program(医师科学家培训计划)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11613 2025-06-16 cs.LG cs.AI 79%

Model Organisms for Emergent Misalignment

Edward Turner, Anna Soligo, Mia Taylor, Senthooran Rajamanoharan, Neel Nanda

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20906 2026-08-24 eess.SY cs.MA cs.RO cs.SY 新提交 78%

A Safety-Driven Architectural Framework for Fail-Operational Drone Swarms in Critical Missions

面向关键任务中故障可运行无人机集群的安全驱动架构框架

Luiz Giacomossi, Zafer Yigit, Marwan Shakarna, Shoaib Saleemi, Ivan Tomasic, Baran Çurüklü, Håkan Forsberg

专题命中 其他安全 :safety(title,abstract)

AI总结 针对安全关键操作中无人机集群的认证需求,提出结合SAE ARP4754B方法的混合关键度架构框架,通过硬件隔离的安全监控器实现飞行关键核心与集群管理器解耦,经马尔可夫建模可满足危险故障要求。

Comments 10 pages, 7 figures. Accepted for presentation at the 45th AIAA/IEEE Digital Avionics Systems Conference (DASC), Orlando, FL, USA, 2026. \c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15698 2026-08-24 cs.CV cs.IR 版本更新 78%

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

ConceptFormer:学习自适应潜在概念以实现视觉文档检索中的查询-文档对齐

Chunyi Peng, Zhipeng Xu, Yukun Yan, Zhenghao Liu, Shi Yu, Sen Mei, Yubo Sun, Yongheng Zhang, Jie Zhou, Yu Gu, Ge Yu, Maosong Sun

机构 * Northeastern University(东北大学) Tsinghua University(清华大学) Peking University(北京大学)

专题命中 其他安全 :alignment(title,abstract)

AI总结 针对视觉文档检索中现有监督信号的局限,本文提出ConceptFormer框架,以自适应潜在概念为中间表示衔接语义鸿沟,在基准测试中较最强基线实现了16.7%、22.1%的NDCG@10相对提升,性能优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21948 2026-08-24 cs.CV 版本更新 78%

Deep Models, Shallow Alignment: Uncovering the Granularity Mismatch in Neural Decoding

深度模型,浅层对齐:揭示神经解码中的粒度不匹配

Yang Du, Siyuan Dai, Yonghao Song, Paul M. Thompson, Haoteng Tang, Liang Zhan

机构 * Dept. of Electrical & Computer Engineering, University of Pittsburgh, USA(宾夕法尼亚大学电气与计算机工程系) Dept. of Biomedical Engineering, Tsinghua University, China(清华大学生物医学工程系) Dept. of Neurology, University of Southern California, USA(美国南加州大学神经病学系) Dept. of Computer Science, University of Texas Rio Grande Valley, USA(德克萨斯理工大学里奥格兰德谷分校计算机科学系)

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文提出浅层对齐方法,通过对比学习策略解决神经解码中的粒度不匹配问题,显著提升解码性能。

Comments 33 pages, 16 figures

Journal ref Transactions on Machine Learning Research (2026), ISSN 2835-8856

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18118 2026-08-20 math.OC cs.SY eess.SY 新提交 78%

Formal Safety Verification for Nonlinear Systems with Generative Barrier Certificate

基于生成式障碍证书的非线性系统形式化安全验证

Mengxin Ren, Hanrui Zhao

专题命中 其他安全 :safety(title,abstract)

AI总结 该研究针对非线性系统安全验证中障碍证书推导计算成本高的问题,提出基于大语言模型的生成式框架,将BMI问题转化为LMI测试,实现了远超传统方法的速度与性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏