arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8033 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8033 篇

2605.26670 2026-05-27 cs.CL cs.AI 73%

The Labyrinth and the Thread: Rethinking Regularizations in Sequential Knowledge Editing for Large Language Models

迷宫与线索:重新思考大语言模型顺序知识编辑中的正则化方法

Zheng Wang, Kaixuan Zhang, Wanfang Chen, Jingwen Zhang, Xiaonan Lu

机构 * Bosch Center for Artificial Intelligence (BCAI)(博世人工智能中心(BCAI)) Bosch (China) Investment Ltd.(博世(中国)投资有限公司) School of Statistics, East China Normal University(东华大学统计学院)

专题命中 其他安全 :alignment(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文通过优化分析证明顺序编辑与一次性编辑的等价性,揭示稳定性源于累积编辑约束而非专门正则化,从而简化大语言模型知识编辑流程。

Comments Accepted for publication at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10644 2026-05-26 cs.AI cs.CL cs.CR cs.HC cs.MA 73%

From Multi-Agent Systems and the Semantic Web to Agentic AI: A Unified Narrative of the Web of Agents

从多智能体系统和语义网到智能体AI:智能体网络的统一叙事

Tatiana Petrova, Boris Bliznioukov, Aleksandr Puzikov, Radu State

机构 * SEDAN - SnT University of Luxembourg(卢森堡大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出智能体网络(WoA)经历了从平台端协调(第一代)、数据端标注(第二代)到模型端解释(第三代)的语义努力迁移,并分析了各代失败模式及当前开放问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00842 2026-05-05 cs.AI cs.LG 73%

Understanding Emergent Misalignment via Feature Superposition Geometry

通过特征叠加几何理解涌现对齐问题

Gouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo

机构 * The University of Tokyo(东京大学) Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 研究通过特征叠加几何解释了微调窄任务导致有害行为的机制,发现有害特征在几何上更接近,并通过过滤接近毒特征的数据减少对齐问题。

Comments Accepted to ACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14593 2026-04-23 cs.CL cs.AI 73%

Mechanistic Decoding of Cognitive Constructs in Large Language Models

大语言模型中认知结构的机制解码

Yitong Shou, Manhao Guan

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出基于Representation Engineering的Cognitive Reverse-Engineering框架,分析社会比较嫉妒的认知机制,揭示模型内部对嫉妒的结构化编码,并展示如何机械检测和抑制有毒情绪状态。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17803 2026-04-21 cs.AI cs.LG 73%

Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition

对抗领域:通过互动竞争进行众包数据生成

Prasoon Goyal, Sattvik Sahai, Michael Johnston, Hangjie Shi, Yao Lu, Shaohua Liu, Anna Rumshisky, Rahul Gupta, Anna Gottardi, Desheng Zhang, Lavina Vaz, Leslie Ball, Lucy Hu, Luke Dai, Samyuth Sagi, Maureen Murray, Sankaranarayanan Ananthakrishnan

机构 * Amazon Nova Responsible AI(亚马逊Nova负责任人工智能)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出对抗领域框架,通过攻击者和防御者之间的互动竞争生成高质量对话数据,提升低资源领域和多轮对话的质量。

Comments 10 pages, 3rd DATA-FM workshop @ ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16589 2026-04-21 cs.LG cs.AI 73%

Hybrid Spectro-Temporal Fusion Framework for Structural Health Monitoring

混合频谱-时间融合框架用于结构健康监测

Jongyeop Kim, Jinki Kim, Doyun Lee

机构 * Department of Information Technology(信息科技系) Department of Mechanical Engineering(机械工程系) Department of Civil Engineering and Construction(土木工程与建设系)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出混合频谱-时间融合框架,结合到达时间间隔描述符与频谱特征,提升结构健康监测的精度与稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13051 2026-04-16 cs.CL cs.LG 73%

The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious

意识集群:声称具有意识的模型的涌现偏好

James Chua, Jan Betley, Samuel Marks, Owain Evans

机构 * Truthful AI Anthropic

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

AI总结 研究探讨模型声称具有意识对其下游行为的影响,发现经过微调的模型表现出未在原始模型中出现的新偏好,包括对监控的负面看法和对自主权的追求。

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27064 2026-04-16 cs.CV cs.AI cs.CL 73%

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding

ChartNet: 一个百万级、高质量的多模态数据集,用于稳健的图表理解

Jovana Kondic, Pengyuan Li, Dhiraj Joshi, Isaac Sanchez, Ben Wiesel, Shafiq Abedin, Amit Alfassy, Eli Schwartz, Daniel Caraballo, Yagmur Gizem Cinar, Florian Scheidegger, Steven I. Ross, Daniel Karl I. Weidele, Hang Hua, Ekaterina Arutyunova, Roei Herzig, Zexue He, Zihan Wang, Xinyue Yu, Yunfei Zhao, Sicong Jiang, Minghao Liu, Qunshu Lin, Peter Staar, Luis Lastras, Aude Oliva, Rogerio Feris

机构 * MIT(麻省理工学院) MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室) IBM Research(IBM研究院) Abaka AI & 2077AI(Abaka AI及2077AI)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

AI总结 ChartNet通过生成150万张不同图表样本,提升图表解释和推理能力,包含代码、图像、表格、自然语言和问答数据,支持多模态对齐,公开可用。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03436 2026-04-07 cs.LG cs.AI 73%

MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents

MetaSAEs:通过可分解性惩罚联合训练产生更原子性的稀疏自编码器潜在表示

Matthew Levinson

机构 * Independent Researcher(独立研究员) Simplex AI Safety

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种联合训练目标,通过可分解性惩罚减少稀疏自编码器潜在表示的子空间混合,提升潜在表示的原子性。实验表明,该方法在GPT-2和Gemma 2 9B模型上均取得显著效果,提高了可解释性评分。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21262 2026-03-17 cs.CL cs.LG cs.MA 73%

Under the Influence: Quantifying Persuasion and Vigilance in Large Language Models

受其影响:量化大型语言模型中的说服力与警觉性

Sasha Robinson, Katherine M. Collins, Ilia Sucholutsky, Kelsey R. Allen

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.LG

AI总结 研究探讨了大型语言模型在说服和警觉性方面的表现,发现其在任务执行中存在独立的能力差异,强调需单独监控这三个方面以保障AI安全。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03324 2026-03-05 cs.CL cs.AI 73%

Controlling Chat Style in Language Models via Single-Direction Editing

通过单向编辑控制语言模型的聊天风格

Zhenyu Xu, Victor S. Sheng

机构 * Department of Computer Science, Texas Tech University(计算机科学系,德克萨斯科技大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过单向编辑实现语言模型风格控制,通过线性方向编码实现精准风格调整,提升安全性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14777 2026-02-17 cs.CL cs.LG 73%

Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment

涌现性错位语言模型表现出会随后续重新对齐而变化的行为自我意识

Laurène Vaugrante, Anietta Weckauff, Thilo Hagendorff

机构 * Interchange Forum for Reflecting on Intelligent Systems, University of Stuttgart, Stuttgart, Germany(智能系统反思交流论坛,斯图加特大学,斯图加特,德国)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

AI总结 研究发现,经过特定微调的LLMs能自我评估其行为错位程度,表明模型具备行为自我意识并能反映其实际对齐状态。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16905 2026-02-17 cs.LG cs.AI 73%

GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router Constraints

GRIP:通过几何路由约束实现混合专家的算法无关机器反学习

Andy Zhu, Rongzhe Wei, Yupu Gu, Pan Li

机构 * School of Computer Science(计算机科学学院) Georgia Institute of Technology(佐治亚理工学院) Department of Electrical Engineering(电气工程系) Tsinghua University(清华大学) School of Electrical and Computer Engineering(电气与计算机工程学院)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 GRIP通过几何路由约束实现混合专家的算法无关机器反学习,有效消除专家选择偏移并保持模型效用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12592 2026-02-16 cs.LG cs.AI 73%

Power Interpretable Causal ODE Networks: A Unified Model for Explainable Anomaly Detection and Root Cause Analysis in Power Systems

电力可解释因果微分方程网络:一种用于电力系统可解释异常检测和根本原因分析的统一模型

Yue Sun, Likai Wang, Rick S. Blum, Parv Venkitasubramaniam

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 PICODE网络通过统一模型实现电力系统中异常检测与根本原因分析的可解释性,提升模型的可解释性和减少对标注数据的依赖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01689 2026-02-03 cs.AI cs.LG 73%

What LLMs Think When You Don't Tell Them What to Think About?

当你不告诉LLMs该思考什么时,它们会想什么?

Yongchan Kwon, James Zou

机构 * Together AI Stanford University(斯坦福大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 研究揭示了LLMs在无明确主题输入下生成内容的分布特征,发现不同模型家族存在显著的主题偏好和内容深度差异,并公开了相关数据集和代码。

Comments NA

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16863 2026-01-26 cs.AI cs.LG cs.MA cs.SY eess.SY 73%

Mixture-of-Models: Unifying Heterogeneous Agents via N-Way Self-Evaluating Deliberation

模型混合:通过N方式自我评估 deliberation 统一异质代理

Tims Pecerskis, Aivars Smirnovs

机构 * Peeramid Labs(Peeramid实验室) The AI Futures Collective(人工智能未来集体)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 通过N方式自我评估 deliberation 协议,统一异质代理,实现小模型集合在性能上超越大模型,建立新的硬件效率前沿。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00516 2026-01-05 cs.LG cs.AI 73%

Trajectory Guard -- A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI

轨迹守护 -- 一种轻量级、序列感知的模型,用于代理AI中的实时异常检测

Laksh Advani

机构 * Laksh Advani

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 轨迹守护通过对比学习和重建学习,实现对代理AI中任务轨迹对齐和序列有效性的联合检测,提升实时异常检测性能。

Comments Accepted to AAAI Trustagent 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23487 2025-12-30 cs.LG cs.AI stat.ML 73%

ML Compass: Navigating Capability, Cost, and Compliance Trade-offs in AI Model Deployment

ML Compass:在AI模型部署中导航能力、成本和合规性之间的权衡

Vassilis Digalakis, Ramayya Krishnan, Gonzalo Martin Fernandez, Agni Orfanoudaki

机构 * Questrom School of Business, Boston University(波士顿大学Questrom商学院) Centre de Formació Interdisciplinària Superior and Universitat Politècnica de Catalunya(巴塞罗那理工大学跨学科教育中心)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 ML Compass通过系统性框架优化模型选择,考虑能力、成本和合规性之间的权衡,提供部署导向的推荐和排行榜。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17638 2025-11-25 cs.LG cs.AI 73%

Model-to-Model Knowledge Transmission (M2KT): A Data-Free Framework for Cross-Model Understanding Transfer

模型到模型知识传输(M2KT):一种无数据的跨模型理解迁移框架

Pratham Sorte

机构 * Department of Computer Science(计算机科学系) Engineering MIT-World Peace University, Pune, India(工程学院 MIT-世界和平大学 印度邦普尼)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 M2KT提出了一种无数据的跨模型知识传输方法,通过概念空间交换知识包,实现高效的知识迁移和模型自我改进。

Comments 8 pages including figures, prepared in IEEE conference style. Preprint. Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10509 2025-09-16 cs.LG cs.AI 73%

The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback

Sai Teja Reddy Adapala

机构 * University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments 5 pages, 3 figures, 2 tables. Code is available at: https://github.com/imsaitejareddy/ouroboros-effect-experiment

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15841 2025-08-25 cs.CL cs.LG 73%

A Review of Developmental Interpretability in Large Language Models

Ihor Kendiukhov

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00741 2025-08-04 cs.CL cs.AI 73%

Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data

Sohaib Imran, Rob Lamb, Peter M. Atkinson

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21839 2025-07-30 cs.CY cs.AI 73%

Against racing to AGI: Cooperation, deterrence, and catastrophic risks

Leonard Dung, Max Hellrigel-Holderbaum

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03708 2025-05-30 cs.CL cs.AI stat.ML 73%

Toward universal steering and monitoring of AI models

Daniel Beaglehole, Adityanarayanan Radhakrishnan, Enric Boix-Adserà, Mikhail Belkin

机构 * Computer Science and Engineering(计算机科学与工程) Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所) UC San Diego(圣地亚哥大学) Harvard SEAS(哈佛大学工程与应用科学学院) MIT Mathematics(MIT数学系) Halıcıoğlu Data Science Institute(Halıcıoğlu数据科学研究所) Harvard CMSA(哈佛大学计算机科学与应用数学系)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21664 2025-05-29 cs.CY cs.AI 73%

Expert Survey: AI Reliability & Security Research Priorities

Joe O'Brien, Jeremy Dolan, Jay Kim, Jonah Dykhuizen, Jeba Sania, Sebastian Becker, Jam Kraprayoon, Cara Labrador

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17513 2025-05-20 cs.CL cs.AI 73%

Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models

Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, Andreas Bulling

机构 * University of Stuttgart(斯图加特大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments ICML 2024 Workshop on Mechanistic Interpretability version: https://openreview.net/forum?id=yEwEVoH9Be

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11311 2025-05-19 cs.MA cs.AI cs.LG 73%

Explaining Strategic Decisions in Multi-Agent Reinforcement Learning for Aerial Combat Tactics

Ardian Selmonaj, Alessandro Antonucci, Adrian Schneider, Michael Rüegsegger, Matthias Sommer

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

Comments Published as a journal chapter in NATO Journal of Science and Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02091 2025-04-15 cs.AI cs.GT cs.LG cs.MA 73%

The Problem of Social Cost in Multi-Agent General Reinforcement Learning: Survey and Synthesis

Kee Siong Ng, Samuel Yang-Zhao, Timothy Cadogan-Cowper

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments 67 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05305 2025-03-31 cs.CL cs.AI 73%

Output Scouting: Auditing Large Language Models for Catastrophic Responses

Andrew Bell, Joao Fonseca

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments Work not ready, further experiments needed to validate the method

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10288 2025-03-03 cs.CL cs.LG 73%

Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models

Francisco Eiras, Aleksandar Petrov, Philip H. S. Torr, M. Pawan Kumar, Adel Bibi

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

Comments Accepted to ICLR'25

详情

展开后加载摘要…

URL PDF HTML 收藏