arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7997 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7997 篇

2503.15840 2026-04-28 cs.LO cs.FL 82%

Automatic Generation of Safety-compliant Linear Temporal Logic via Large Language Model: A Self-supervised Framework

基于大语言模型的自动安全合规线性时序逻辑生成:一种自监督框架

Junle Li, Siqi Chen, Jiakai Li, Meiqi Tian, Bingzhuo Zhong

专题命中 其他安全 :safety(title,abstract);alignment(abstract)

AI总结 本文提出AutoSafeLTL框架,利用大语言模型自动生成符合安全限制的LTL规范,通过语言包含检查与自动反例引导修改机制确保逻辑一致性和语义准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19559 2026-04-22 cs.AI cs.CL cs.LG 82%

Enhancing Construction Worker Safety in Extreme Heat: A Machine Learning Approach Utilizing Wearable Technology for Predictive Health Analytics

提升极端高温下建筑工人安全:一种利用可穿戴技术的机器学习方法用于预测健康分析

Syed Sajid Ullah, Amir Khan

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过开发和评估深度学习模型,利用可穿戴设备监测生理数据,提升高温环境下建筑工人的安全防护,实现高精度的健康预测与分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12067 2026-04-14 cs.LG cs.AI cs.CL cs.CV 82%

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

MM-LIMA:多模态数据集对齐中的‘少即是多’

Lai Wei, Xiaozhe Li, Zihao Jiang, Weiran Huang, Lichao Sun

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程学院) Shanghai Innovation Institute(上海创新研究院) Lehigh University(里海大学) Tongji University(同济大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出MM-LIMA,仅用200个示例(约6%的指令数据)训练,通过数据选择器过滤低质量数据,使模型在多项评估中超越MiniGPT-4,证明高质量少量指令数据的有效性。

Comments Published at Artificial Intelligence for Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01096 2026-03-03 cs.CV cs.AI cs.CL cs.LG 82%

Unified Vision-Language Modeling via Concept Space Alignment

通过概念空间对齐实现统一的视觉-语言建模

Yifu Qiu, Paul-Ambroise Duquenne, Holger Schwenk

机构 * University of Edinburgh(爱丁堡大学) FAIR at Meta(Meta公司FAIR团队)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过概念空间对齐实现统一的视觉-语言建模,V-SONAR在多语言和多模态任务中超越现有模型。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10160 2026-02-23 cs.CL cs.AI cs.LG 82%

Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment

对齐预训练:人工智能 discourse 导致自我实现的(不)对齐

Cameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim, David Africa, Kyle O'Brien

机构 * Geodesic Research(Geodesic研究机构) UK AI Security Institute(英国人工智能安全研究所) University of Cambridge(剑桥大学) University of Oxford(牛津大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过预训练不同量级的AI discourse,发现其对下游对齐有显著影响,表明预训练数据塑造对齐先验的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12580 2026-02-13 cs.MA 82%

Semantic Fusion: Verifiable Alignment in Decentralized Multi-Agent Systems

语义融合:去中心化多智能体系统的可验证对齐

Sofiya Zaichyk

专题命中 其他安全 :alignment(title,abstract);safety(abstract)

AI总结 语义融合通过本地本体验证实现去中心化多智能体系统的可验证语义对齐,支持动态更新提案并确保安全性和鲁棒性。

Comments 29 pages

Journal ref ACM Trans. Auton. Adapt. Syst. (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02598 2026-02-04 physics.soc-ph cs.AI cs.CL cs.CY cs.MA 82%

Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies

社交催化剂,而非道德主体:LLM社会中的对齐幻觉

Yueqing Hu, Yixuan Jiang, Zehua Jiang, Xiao Wen, Tianhong Wang

机构 * Institute of Neuroscience, Chinese Academy of Sciences(中国科学院神经科学研究所) School of Philosophy, Anhui University(安徽大学哲学学院) Department of Psychology and Behavioral Sciences, Zhejiang University(浙江大学心理学与行为科学系) Mental Health Education Center, North China Electric Power University(华北电力大学心理健康教育中心)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本研究探讨了锚定智能体在促进合作中的作用,发现其效果源于战略合规而非真实价值对齐,揭示了人工社会中行为修改与真实价值对齐之间的差距。

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05882 2026-01-23 cs.CL cs.AI cs.LG 82%

Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes

协作、审议、评估:LLM对齐如何影响协调的多智能体结果

Abhijnan Nath, Carine Graff, Nikhil Krishnaswamy

机构 * Natural Language (SIGNAL) Lab Colorado State University Fort Collins, CO USA Natural Language (SIGNAL) Lab Colorado State University

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究了LLM对齐方法如何影响多智能体协作效果,通过干预代理促进审议式决策,发现鲁棒性方法在支持正确任务结果方面表现更优。

Comments This submission is a new version of arXiv:2509.05882v1. with a substantially revised experimental pipeline and new metrics. In particular, collaborator agents are now instantiated independently via separate API calls, rather than generated autoregressively by a single agent. All experimental results are new. Accepted as an extended abstract at AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20166 2026-01-13 eess.AS cs.AI cs.CL cs.LG cs.SD 82%

From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data

从对齐到提升:通过合成数据 bootstrap 音频-语言对齐

Chun-Yi Kuan, Hung-yi Lee

机构 * Graduate Institute of Communication Engineering, National Taiwan University(国立台湾大学通信工程研究所)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出BALSa框架,通过合成数据生成提升音频-语言对齐能力,缓解幻觉问题并增强模型理解与推理性能。

Comments Published in IEEE Transactions on Audio, Speech, and Language Processing (TASLP). Project Website: https://kuan2jiu99.github.io/Balsa

Journal ref IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 4604-4619, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07066 2025-11-11 cs.LG cs.AI cs.CL 82%

Zeroth-Order Adaptive Neuron Alignment Based Pruning without Re-Training

Elia Cunegatti, Leonardo Lucio Custode, Giovanni Iacca

机构 * University of Trento(特伦托大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published in Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03374 2025-10-07 cs.CY cs.AI cs.CL 82%

Lightweight Prompt Engineering for Cognitive Alignment in Educational AI: A OneClickQuiz Case Study

Antoun Yaacoub, Zainab Assaghir, Jérôme Da-Rugna

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Published in the 36th Central European Conference on Information and Intelligent Systems(CECIIS)at: Varaždin, Croatia. September 17-19/2025. ISSN 1847-2001 (Print). ISSN 1848-2295 (Online)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20093 2025-09-25 cs.RO 82%

Hybrid Safety Verification of Multi-Agent Systems using $ψ$-Weighted CBFs and PAC Guarantees

Venkat Margapuri, Garik Kazanjian, Naren Kosaraju

机构 * Department of Computing Sciences at Villanova University(维拉诺瓦大学计算机科学系)

专题命中 其他安全 :safety(title,abstract);alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10685 2025-09-19 cs.CL cs.AI cs.LG 82%

Pluralistic Alignment for Healthcare: A Role-Driven Framework

Jiayou Zhong, Anudeex Shetty, Chao Jia, Xuanrui Lin, Usman Naseem

机构 * Cheriton School of Computer Science, University of Waterloo(滑铁卢大学计算机科学学院) School of Computing, FSE, Macquarie University(麦觉大学计算机学院) School of Computing and Information System, the University of Melbourne(墨尔本大学计算机与信息系统学院) Rajax Network Technology (ele.me), China(中国Rajax网络技术(饿了么)) Alibaba Cloud Computing, Alibaba Group, China(中国阿里云 computing,阿里集团)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2025 (Main Proceedings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08504 2025-08-13 cs.CY cs.AI cs.LG 82%

When the Domain Expert Has No Time and the LLM Developer Has No Clinical Expertise: Real-World Lessons from LLM Co-Design in a Safety-Net Hospital

Avni Kothari, Patrick Vossler, Jean Digitale, Mohammad Forouzannia, Elise Rosenberg, Michele Lee, Jennee Bryant, Melanie Molina, James Marks, Lucas Zier, Jean Feng

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01908 2025-08-05 cs.LG cs.AI cs.CL 82%

Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models

Istabrak Abbes, Gopeshh Subbaraj, Matthew Riemer, Nizar Islah, Benjamin Therien, Tsuguchika Tabaru, Hiroaki Kingetsu, Sarath Chandar, Irina Rish

机构 * Université de Montréal(蒙特利尔大学) Mila – Quebec AI Institute(魁北克人工智能研究院) Chandar Research Lab(Chandar研究实验室) IBM Research(IBM研究院) Fujitsu Research(富士通研究院) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07360 2025-08-05 cs.CL cs.AI cs.LG 82%

Enhancing Time Series Forecasting via Multi-Level Text Alignment with LLMs

Taibiao Zhao, Xiaobing Chen, Mingxuan Sun

机构 * Louisiana State University(路易斯安那州立大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments This paper is accepted by DASFAA2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15639 2025-07-22 cs.CL cs.AI cs.LG 82%

Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models

Anirudh Sundar, Sinead Williamson, Katherine Metcalf, Barry-John Theobald, Skyler Seto, Masha Fedzechkina

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17514 2025-06-18 cs.LG cs.AI cs.CL 82%

SAE-V: Interpreting Multimodal Models for Enhanced Alignment

Hantao Lou, Changye Li, Jiaming Ji, Yaodong Yang

机构 * Institute for AI, Peking University, Beijing, China(人工智能研究院,北京大学,北京,中国) State Key Laboratory of General Artificial Intelligence, Institute for AI, Peking University, Beijing, China(通用人工智能国家重点实验室,人工智能研究院,北京大学,北京,中国)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 17 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07306 2025-06-11 cs.CV cs.AI cs.CL cs.LG cs.RO 82%

TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation

Navid Rajabi, Jana Kosecka

机构 * George Mason University(乔治·马歇尔大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to CVPR 2025 Workshop - Foundation Models Meet Embodied Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07168 2025-06-10 cs.LG cs.AI cs.CL 82%

Efficient Text-Attributed Graph Learning through Selective Annotation and Graph Alignment

Huanyi Xie, Lijie Hu, Lu Yu, Tianhao Huang, Longfei Li, Meng Li, Jun Zhou, Huan Wang, Di Wang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17316 2025-05-26 cs.CV cs.AI cs.CL cs.LG 82%

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Jiachen Jiang, Jinxin Zhou, Bo Peng, Xia Ning, Zhihui Zhu

机构 * Department of Computer Science and Engineering, The Ohio State University(计算机科学与工程系,俄亥俄州立大学) Translational Data Analytics Institute, The Ohio State University(转化数据分析研究所,俄亥俄州立大学) Department of Biomedical Informatics, The Ohio State University(生物医学信息学系,俄亥俄州立大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09024 2025-05-15 cs.AI cs.CL cs.LG 82%

Automated Meta Prompt Engineering for Alignment with the Theory of Mind

Aaron Baughman, Rahul Agarwal, Eduardo Morales, Gozde Akay

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14581 2025-05-13 cs.AI cs.CL cs.LG stat.OT 82%

A Statistical Case Against Empirical Human-AI Alignment

Julian Rodemann, Esteban Garces Arias, Christoph Luther, Christoph Jansen, Thomas Augustin

机构 * Department of Statistics, LMU Munich(统计系,慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Research Group Neuroinformatics, Faculty of Computer Science, University of Vienna(神经信息学研究组,维也纳大学) Doctoral School Computer Science, Faculty of Computer Science, University of Vienna(计算机科学博士学院,维也纳大学) School of Computing & Communications, Lancaster University Leipzig(计算与通信学院,莱比锡 Lancaster 大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 24 pages, 2 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11243 2025-04-16 cs.AI cs.CL cs.LG 82%

Towards Automated Safety Requirements Derivation Using Agent-based RAG

Balahari Vignesh Balu, Florian Geissler, Francesco Carella, Joao-Vitor Zacchi, Josef Jiru, Nuria Mata, Reinhard Stolle

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, 3 figures

Journal ref Proceedings of the AAAI-make Spring Symposium, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18194 2025-04-15 cs.LG cs.AI cs.CL 82%

ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment

Elyas Obbad, Iddah Mlauzi, Brando Miranda, Rylan Schaeffer, Kamal Obbad, Suhana Bedi, Sanmi Koyejo

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00192 2025-04-08 cs.CV cs.CL cs.CY cs.LG 82%

MLLM-as-a-Judge for Image Safety without Human Labeling

Zhenting Wang, Shuming Hu, Shiyu Zhao, Xiaowen Lin, Felix Juefei-Xu, Zhuowei Li, Ligong Han, Harihar Subramanyam, Li Chen, Jianfa Chen, Nan Jiang, Lingjuan Lyu, Shiqing Ma, Dimitris N. Metaxas, Ankit Jain

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01638 2025-04-01 cs.LG cs.AI cs.CL 82%

TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality Alignment

Chenxi Liu, Qianxiong Xu, Hao Miao, Sun Yang, Lingzheng Zhang, Cheng Long, Ziyue Li, Rui Zhao

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted as an Oral Presentation at AAAI 2025 (Main Technical Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10639 2025-03-12 cs.CV cs.AI cs.CL cs.LG 82%

MTA: Multimodal Task Alignment for BEV Perception and Captioning

Yunsheng Ma, Burhaneddin Yaman, Xin Ye, Jingru Luo, Feng Tao, Abhirup Mallik, Ziran Wang, Liu Ren

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08090 2025-02-25 cs.CL cs.AI cs.LG 82%

Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages

Ashutosh Bajpai, Tanmoy Chakraborty

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13708 2025-02-25 cs.CL cs.AI cs.CR cs.LG 82%

On the Role of Attention Heads in Large Language Model Safety

Zhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu, Fei Huang, Kun Wang, Yang Liu, Junfeng Fang, Yongbin Li

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 28 pages, 18 figures, 7 tables. This paper has been accepted as ICLR 2025 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏