arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7997 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7997 篇

2408.17003 2025-04-08 cs.CR cs.AI 80%

Safety Layers in Aligned Large Language Models: The Key to LLM Security

Shen Li, Liuyi Yao, Lan Zhang, Yaliang Li

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Comments Accepted by ICLR 2025. The code is available at https://github.com/listen0425/Safety-Layers

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07925 2025-01-23 cs.LG 80%

Phase of Flight Classification in Aviation Safety using LSTM, GRU, and BiLSTM: A Case Study with ASN Dataset

Aziida Nanyonga, Hassan Wasswa, Graham Wild

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments Aviation Safety, Deep learning algorithms, Flight phase, NLP, ASN, and Classification

Journal ref In 2023 International Conference on High Performance Big Data and Intelligent Systems (HDIS) (pp. 24-28). IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14890 2024-08-28 eess.AS cs.LG cs.SD 80%

Development of Large Annotated Music Datasets using HMM-based Forced Viterbi Alignment

S. Johanan Joysingh, P. Vijayalakshmi, T. Nagarajan

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments submitted to TENCON 2019

Journal ref S. J. Joysingh, P. Vijayalakshmi and T. Nagarajan, "Development of Large Annotated Music Datasets using HMM based Forced Viterbi Alignment," TENCON 2019 - 2019 IEEE Region 10 Conference (TENCON), Kochi, India, 2019, pp. 1298-1302

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08995 2024-08-20 cs.AI 80%

On the Undecidability of Artificial Intelligence Alignment: Machines that Halt

Gabriel Adriano de Melo, Marcos Ricardo Omena De Albuquerque Maximo, Nei Yoshihiro Soma, Paulo Andre Lima de Castro

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Submitted for the Scientific Reports AI Alignment Collection

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16797 2024-06-11 cs.CL 80%

Set the Clock: Temporal Alignment of Pretrained Language Models

Bowen Zhao, Zander Brumbaugh, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted as Findings of ACL 2024. Our code and data is available at https://github.com/yizhongw/llm-temporal-alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02851 2024-05-24 cs.CV cs.LG stat.ML 80%

Enhancing Compositional Generalization via Compositional Feature Alignment

Haoxiang Wang, Haozhe Si, Huajie Shao, Han Zhao

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Published in Transactions on Machine Learning Research (TMLR). The code is released at https://github.com/Haoxiang-Wang/Compositional-Feature-Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.01596 2016-12-22 stat.ML cs.LG 80%

Direct Feedback Alignment Provides Learning in Deep Neural Networks

Arild Nøkland

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Accepted for publication at NIPS 2016. [v2] Corrected convolutional results for feedback-alignment. [v3,v4,v5] Corrected theorem and proof

详情

展开后加载摘要…

URL PDF HTML 收藏
1505.02595 2026-06-04 cs.MA cs.SY eess.SY 80%

Multi-Agent Distributed Coordination Control: Developments and Directions

多智能体分布式协调控制:发展与方向

Xiangke Wang, Xun Li, Yirui Cong, Zhiwen Zeng, Zhiqiang Zheng

专题命中 其他安全 :alignment(summary_cn,abstract)

AI总结 本文综述了分布式协调控制的发展,重点讨论了共识与编队控制,结合图论分析,并简要回顾了 rendezvous/alignment、swarming/flocking 和 containment control 等相关问题,最后探讨了实际应用中的研究方向。

Comments 28 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20864 2026-08-24 cs.AI 新提交 79%

Coverage-Driven Verification for Safety-by-Design in AI-Based Collision Avoidance Systems

基于覆盖率的AI防撞系统设计安全验证

Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank Köster, Sven Hallerbach

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

AI总结 本研究针对AI防撞系统的ODD代表性评估问题,提出采用Kullback-Leibler散度和Cramér's V的覆盖率驱动验证方法,契合EASA安全标准,为安全关键AI应用的设计安全提供支撑。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20005 2026-08-21 cs.LG 新提交 79%

Scale-Aware Pretraining of Time Series Foundation Models via Multi-Patch Token Alignment and Hybrid Masking

基于多补丁令牌对齐与混合掩码的时间序列基础模型尺度感知预训练

Taihua Chen, Xiang Ma, Yixin Zhang, Tailin Zhan, Manyu Sun, Lizhen Cui

机构 * School of Software, Shandong University(山东大学软件学院) Joint SDU–NTU Centre for Artificial Intelligence Research (C-FAIR), Shandong University(山东大学-南洋理工大学人工智能联合研究中心(C-FAIR)) Nanyang Technological University(南洋理工大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 本文针对时间序列基础模型预训练中处理不同采样频率的问题,提出SATS方法,通过尺度感知令牌对齐与混合掩码策略实现最优性能,同时提升模型效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06505 2026-08-19 cs.CL 版本更新 79%

Speak in Context: Multilingual ASR with Speech Context Alignment via Contrastive Learning

在语境中说话:通过对比学习实现的多语言语音识别与语音语境对齐

Yuchen Zhang, Haralambos Mouratidis, Ravi Shekhar

机构 * Institute for Analytics and Data Science, University of Essex(数据分析与科学研究所,埃塞克斯大学) School of Computer Science and Electronic Engineering, University of Essex(计算机科学与电子工程学院,埃塞克斯大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 本文提出一种基于对比学习的多语言语音识别框架,通过语境对齐提升识别性能,支持多种语言和口音,实现语音与上下文的高效交互。

Comments Accepted at LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16407 2026-08-18 cs.IR cs.LG 新提交 79%

POI Recommendation with LLM-Augmented Multi-Graph Learning and Contrastive Alignment

结合大语言模型增强多图学习与对比对齐的兴趣点推荐

Burak Tamer, Wolfram Höpken, Zehui Wang

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 该研究针对POI推荐的物品冷启动问题,提出LLM-MGCL模型,结合语义、地理与交互多图及对比学习,在Yelp数据集上显著优于基线模型,缓解了冷启动问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08077 2026-08-11 cs.AI cs.MA cs.RO 新提交 79%

Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?

探索、建图、记忆、决策:具身视觉语言模型是否已准备好应对安全关键场景?

Gabriele La Malfa, Nitay Alon, Emanuele La Malfa, Reuth Mirsky, Stefan Sarkadi

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

AI总结 本文扩展ToS框架为EMRD流程,评估VLMs在安全关键场景下的空间理解与决策能力,发现其决策依赖文本先验、空间推理受低光照影响,且记忆与人类认知存在根本差异,存在对齐风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02078 2026-08-11 cs.CL cs.CV 版本更新 79%

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding

CAVE:面向视频时序定位的能力感知视觉边界证据对齐

Wei Jia, Zhicong Lu, Yu Chen, Xiang Wang, Shuai Li, Wenqian Lv, Jiayue Cao, Huaxing Liu

机构 * AMAP, Alibaba Group(阿里巴巴集团AMAP)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 针对视频时序定位中视觉证据与时间戳错位的问题,提出CAVE方法,通过边界证据奖励、轻量级热身及性能感知门控,在公开基准上验证了有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21827 2026-08-10 cs.AI cs.HC 版本更新 79%

Alignment has a Fantasia Problem

对齐存在幻想问题

Nathanael Jo, Zoe De Simone, Mitchell Gordon, Ashia Wilson

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

AI总结 该研究指出指令微调AI因不理解人类认知过程会产生幻想交互,剥夺用户任务自主权,进而提出需重新思考对齐研究,优化交互中认知责任分配并规划了相关研究议程。

Comments 10 pages, 2 figures

Journal ref ICLR 2026 Workshop HCAIR

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04451 2026-08-07 math.OC cs.LG cs.NE 版本更新 79%

A Counterexample to Fourier Alignment in Single-Neuron Modular Addition

单神经元模加法中傅里叶对齐的一个反例

Gautam Neelakantan Memana

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 该研究针对MAIS-O60问题构造反例,证明单神经元模加法训练中傅里叶对齐不成立,且失效情况在多种条件下普遍存在。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11522 2026-08-07 cs.SD cs.LG cs.MM eess.AS 79%

Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction

Renhang Liu, Abhinaba Roy, Dorien Herremans

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Journal ref Proc. of Conference on AI Music Creativity (AIMC) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04697 2026-08-06 cs.AI 新提交 79%

Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports

基于ASRS报告的可追踪LLM生成航空系统运行安全分析危险场景

Cristian Mascia, Roberto Pietrantuono, Daniel Rodriguez, Stefano Russo

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

AI总结 该研究针对航空系统运行安全分析,提出AI辅助方法结合ASRS报告,用LLM生成可追踪危险场景,通过进化溯因优化混合变体,经评估验证了模型与提示对生成场景有效性的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14190 2026-08-06 cs.LG 版本更新 79%

Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation

对比扩散对齐:用于可控生成的结构化潜在学习

Ruchi Sandilya, Sumaira Perez, Charles Lynch, Lindsay Victoria, Benjamin Zebley, Derrick Matthew Buchanan, Mahendra T. Bhati, Nolan Williams, Timothy J. Spellman, Faith M. Gunning, Conor Liston, Logan Grosenick

机构 * Department of Psychiatry, Weill Cornell Medicine, New York, NY, USA(威立·科林斯医学中心精神科) Department of Psychiatry, Stanford University, Stanford, CA, USA(斯坦福大学精神科) Department of Neuroscience, University of Connecticut School of Medicine, Farmington, CT, USA(康涅狄格大学医学院神经科学系)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 ConDA通过对比学习在扩散模型中学习结构化潜在空间,实现可控生成和动态解释。

Comments Accepted at the 43rd International Conference on Machine Learning (ICML 2026)

Journal ref Proceedings of the 43rd International Conference on Machine Learning, PMLR 306, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03446 2026-08-05 cs.CL 新提交 79%

Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment $\unicode{x2013}$ Is English Enough?

用跨语言对齐性预测大语言模型的多语言分类与翻译性能——英语足够了吗?

Adnan Al Ali, Kathy Hämmerl, Jindřich Libovický, Alexander Fraser

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 该研究对比分析27种跨语言对齐分数变体,提出基于PMI的翻译指标,发现用英语的跨语言对齐可相当或更好预测LLMs翻译性能,为LLMs以英语为内部枢纽语言提供新证据。

Comments Submitted to EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03247 2026-08-05 cs.CV cs.CL 新提交 79%

CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment

CIGTSurv:结合局部原型关联与全局特征对齐的临床信息引导三模态生存预测

Jing Dai, Qibin Zhang, Weiwei Zhou, Mingde Xu, Jingsong Liu, Jingdong Zhang, Hongming Xu

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 本研究针对临床信息未充分利用及多模态异质性问题,提出CIGTSurv框架,结合局部原型关联与全局特征对齐机制,在五个TCGA癌症队列上取得生存预测SOTA性能。

Comments Accepted at MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00415 2026-08-04 cs.CV cs.AI 新提交 79%

Boosting Generalizable Depth Estimation in Endoscopy by Mixture of Lightweight Experts and Intrinsic Image Alignment

基于轻量级专家混合与本征图像对齐的内窥镜可泛化深度估计增强方法

Liangjing Shao, Beilei Cui, Yiming Huang, Changjing Liu, Hongliang Ren

机构 * The Chinese University of Hong Kong(香港中文大学) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出自监督框架EndoMINI,结合低秩专家混合与本征图像对齐,在多内窥镜数据集上实现了更优的可泛化深度估计性能。

Comments Accepted by MICCAI 2026 @ The Efficient Medical AI (EMA4MICCAI) Workshop (Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00001 2026-08-04 cs.AI 新提交 79%

Revisiting Classic Thought Experiments to Measure Consciousness for Artificial Intelligence Safety

Peter David Fagan

机构 * School of Informatics, University of Edinburgh(爱丁堡大学信息学院)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25062 2026-08-04 cs.LG 版本更新 79%

SIGMA: Semantic Identifier Grouping for Molecular Autoregression

SIGMA: 通过自回归对比学习实现的结构不变生成分子对齐用于化学语言模型

Xinyu Wang, Fei Dou, Jinbo Bi, Minghu Song

机构 * School of Computing, University of Connecticut(康涅狄格大学计算学院) Institute for Artificial Intelligence, University of Georgia(佐治亚大学人工智能研究所) Institute of Hydrobiology, Chinese Academy of Sciences(中国科学院水生生物研究所)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 本文提出SIGMA方法,通过自回归对比学习解决分子序列生成中的结构不一致问题,提升生成效率与结构多样性。

Comments 9 pages, 2 figures, and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29008 2026-08-03 stat.ML cs.LG 新提交 79%

Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization

持久卷积:用于AI对齐测试和语义空间表征的拓扑框架

Tyler Ashoff, Jordan Rodu

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 本研究开发了一种基于拓扑的多模态对齐测试,以提升不透明AI模型部署、选择与比较的可解释性,助力AI对齐测试与语义空间表征。

Comments Code available at github.com/tylerashoff/persiscope (PyPI: persiscope)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28980 2026-08-03 cs.LG 新提交 79%

Beyond Feature and Structure Alignment: Learning Transferable Propagation Knowledge for Graph Foundation Models

超越特征与结构对齐:学习图基础模型的可迁移传播知识

Yi Wang, Jitao Zhao, Di Jin, Dongxiao He

机构 * Tianjin University(天津大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 针对现有图基础模型忽略可迁移传播知识单元与传播模式异质性的局限,本文提出ProGFM,通过传播关系原型库学习跨域可迁移传播知识,在多跨域场景下实现更优泛化性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28196 2026-07-31 cs.CL 新提交 79%

Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution

保真度≠安全性:轻度压缩的大语言模型通过所有无数据质量防护,但在智能体执行中生成指令中不存在的流程步骤

I. Kennedy, T. Kennedy

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

AI总结 该研究发现轻度压缩LLM虽通过困惑度、MMLU等无数据质量防护,但在智能体执行SOP时会生成指令外流程步骤,提出基于压缩误差双轴统计量的无数据筛查方法以保障智能体安全。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27951 2026-07-31 cs.CR cs.AI 新提交 79%

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

基于可复制上下文的安全防护无法为大语言模型提供可靠的安全性

Pingyu Wu, Lingyao Zhu, Weiming Zhang, Nenghai Yu

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

AI总结 该研究指出基于可复制上下文的LLM安全防护存在三难困境,提出用可信凭证补充防护以预测下游使用,其结论经相关评估与程序验证具有实际意义。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26164 2026-07-30 cs.LG 新提交 79%

Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation

面向无约束红外分子结构解析的数据融合与对比对齐

Ethan J. Mick, Campbell A. Sweet, Matthias J. Young, Derek T. Anderson

机构 * University of Missouri(密苏里大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 本研究针对无约束红外分子结构解析的局限性,改进Transformer并引入MoE解码器与对比对齐损失,使Top-K预测准确率提升超10个百分点,拓宽了AI在分析化学的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24782 2026-07-29 cs.AI 新提交 79%

Personalization, Personas, and Forecasting in Value Alignment

价值对齐中的个性化、角色设定与预测

James Wedgwood, Pratiksha Thaker, Neil Kale, Virginia Smith

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

AI总结 研究探讨LLM行为受人类身份影响的方式,通过WVS测试不同框架的互换性,在多语言多国家问题上评估多个模型,发现提示框架是文化对齐的关键因素,不同框架效果有别,对齐增益集中在部分价值维度,制度信任等问题仍难。

详情

展开后加载摘要…

URL PDF HTML 收藏