arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7997 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7997 篇

2411.12275 2024-11-20 cs.CY cs.AI cs.CL 82%

Building Trust: Foundations of Security, Safety and Transparency in AI

Huzaifa Sidhpurwala, Garth Mollett, Emily Fox, Mark Bestavros, Huamin Chen

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17827 2024-11-12 cs.CV cs.AI cs.CL cs.LG 82%

Unified Lexical Representation for Interpretable Visual-Language Alignment

Yifan Li, Yikai Wang, Yanwei Fu, Dongyu Ru, Zheng Zhang, Tong He

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08922 2024-10-28 cs.CL cs.AI cs.LG 82%

Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring

Hee-Jun Jung, Doyeon Kim, Seung-Hoon Na, Kangil Kim

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments This work has been submitted to the ELSEVIER for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00091 2024-09-04 cs.CL cs.AI cs.LG 82%

Classification of Safety Events at Nuclear Sites using Large Language Models

Mishca de Costa, Muhammad Anwar, Daniel Lau, Issam Hammad

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Journal ref 43rd Annual CNS Conference and the 48th Annual CNS/CNA Student Conference Sheraton Cavalier Saskatoon Hotel, Saskatoon, SK, Canada, June 16-19, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02429 2024-08-27 cs.IR cs.AI cs.CL cs.LG 82%

CALRec: Contrastive Alignment of Generative LLMs for Sequential Recommendation

Yaoyiran Li, Xiang Zhai, Moustafa Alzantot, Keyi Yu, Ivan Vulić, Anna Korhonen, Mohamed Hammad

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments RecSys 2024 (Long Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04614 2024-08-15 cs.CL cs.AI cs.LG 82%

Better Alignment with Instruction Back-and-Forth Translation

Thao Nguyen, Jeffrey Li, Sewoong Oh, Ludwig Schmidt, Jason Weston, Luke Zettlemoyer, Xian Li

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13229 2024-06-21 cs.CL cs.AI cs.LG 82%

Probing the Emergence of Cross-lingual Alignment during LLM Training

Hetong Wang, Pasquale Minervini, Edoardo M. Ponti

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to Findings of the Association for Computational Linguistics: ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14863 2024-05-24 cs.CL cs.AI cs.LG 82%

A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns

Asaf Yehudai, Taelin Karidi, Gabriel Stanovsky, Ariel Goldstein, Omri Abend

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments CogSci

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00688 2024-05-03 cs.RO cs.AI cs.CL cs.HC cs.LG 82%

Understanding Social Perception, Interactions, and Safety Aspects of Sidewalk Delivery Robots Using Sentiment Analysis

Yuchen Du, Tho V. Le

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 34 pages, 7 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09704 2024-03-18 cs.CL cs.AI cs.LG 82%

Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations

Swapnaja Achintalwar, Ioana Baldini, Djallel Bouneffouf, Joan Byamugisha, Maria Chang, Pierre Dognin, Eitan Farchi, Ndivhuwo Makondo, Aleksandra Mojsilovic, Manish Nagireddy, Karthikeyan Natesan Ramamurthy, Inkit Padhi, Orna Raz, Jesus Rios, Prasanna Sattigeri, Moninder Singh, Siphiwe Thwala, Rosario A. Uceda-Sosa, Kush R. Varshney

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02896 2024-02-06 cs.CL cs.AI cs.CY cs.MA 82%

LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models

Ivar Frisch, Mario Giulianelli

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments To appear in Proceedings of the 1st Personalization of Generative AI Workshop, EACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08379 2023-11-29 cs.CY cs.AI cs.LG 82%

Scheming AIs: Will AIs fake alignment during training in order to get power?

Joe Carlsmith

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY、cs.LG

Comments 127 pages, 8 figures. Revised again to correct typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10111 2023-11-20 cs.CV cs.AI cs.CL cs.LG 82%

VideoCon: Robust Video-Language Alignment via Contrast Captions

Hritik Bansal, Yonatan Bitton, Idan Szpektor, Kai-Wei Chang, Aditya Grover

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 22 pages, 19 Figures, 7 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.07172 2023-09-15 cs.AI cs.CL cs.LG 82%

Exploring Large Language Models for Ontology Alignment

Yuan He, Jiaoyan Chen, Hang Dong, Ian Horrocks

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ISWC 2023 (Posters and Demos)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05382 2023-09-07 cs.CL cs.AI cs.LG 82%

ChatGPT is on the Horizon: Could a Large Language Model be Suitable for Intelligent Traffic Safety Research and Applications?

Ou Zheng, Mohamed Abdel-Aty, Dongdong Wang, Zijin Wang, Shengxuan Ding

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Submitted to Nature - Machine Intelligence (Revised and Extended)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04275 2023-08-09 cs.CL cs.AI cs.LG 82%

In-Context Alignment: Chat with Vanilla Language Models Before Fine-Tuning

Xiaochuang Han

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.15952 2022-06-13 cs.CL cs.AI cs.LG 82%

Knowledge Graph - Deep Learning: A Case Study in Question Answering in Aviation Safety Domain

Ankush Agarwal, Raj Gite, Shreya Laddha, Pushpak Bhattacharyya, Satyanarayan Kar, Asif Ekbal, Prabhjit Thind, Rajesh Zele, Ravi Shankar

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments LREC 2022 Main Conference Accepted Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.04948 2020-06-11 cs.CY cs.AI cs.LG 82%

AI Research Considerations for Human Existential Safety (ARCHES)

Andrew Critch, David Krueger

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.00783 2018-04-24 cs.AI cs.CY cs.LG 82%

Brief Notes on Hard Takeoff, Value Alignment, and Coherent Extrapolated Volition

Gopal P. Sarma

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY、cs.LG

Comments 3 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21311 2026-08-05 cs.CV cs.AI cs.LG 版本更新 82%

Efficient unsupervised domain adaptation via self-supervised vision transformer and synergistic cross-domain alignment

基于自监督视觉Transformer与协同跨域对齐的高效无监督域适应

Ali Abedi, Q. M. Jonathan Wu, Ning Zhang, Farhad Pourpanah

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 提出EUDA框架,以冻结的DINOv2为特征提取器,结合SDAL损失,在多数据集上实现高效无监督域适应,可训练参数减少42%至99.7%,适配资源受限环境。

Comments 22 pages, 4 figures

Journal ref Abedi, A., Wu, Q.M.J., Zhang, N. et al. Efficient unsupervised domain adaptation via self-supervised vision transformer and synergistic cross-domain alignment. Int. J. Mach. Learn. & Cyber. 17, 423 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06966 2025-09-10 eess.SP cs.AI cs.LG 82%

Cross-device Zero-shot Label Transfer via Alignment of Time Series Foundation Model Embeddings

Neal G. Ravindra, Arijit Sehanobish

机构 * Independent Researcher(独立研究者)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 5 pages, 3 figures, 1 table. tl;dr: Adversarial alignment of Time-Series Foundation Model (TSFM) embeddings enables transfer of high-quality clinical labels from medical-grade to consumer-grade wearables, enabling zero-shot prediction of gestational age without requiring paired data

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04231 2025-06-03 cs.MA cs.AI cs.CY cs.GT 82%

Quantifying Misalignment Between Agents: Towards a Sociotechnical Understanding of Alignment

Aidan Kierans, Avijit Ghosh, Hananel Hazan, Shiri Dori-Hacohen

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments 7 pages, 8 figures, 3 tables, forthcoming at the AAAI-25 Special Track on AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05282 2025-04-10 cs.CY cs.AI 82%

International Scientific Report on the Safety of Advanced AI (Interim Report)

Yoshua Bengio, Sören Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Danielle Goldfarb, Hoda Heidari, Leila Khalatbari, Shayne Longpre, Vasilios Mavroudis, Mantas Mazeika, Kwan Yee Ng, Chinasa T. Okolo, Deborah Raji, Theodora Skeadas, Florian Tramèr, Bayo Adekanmbi, Paul Christiano, David Dalrymple, Thomas G. Dietterich, Edward Felten, Pascale Fung, Pierre-Olivier Gourinchas, Nick Jennings, Andreas Krause, Percy Liang, Teresa Ludermir, Vidushi Marda, Helen Margetts, John A. McDermid, Arvind Narayanan, Alondra Nelson, Alice Oh, Gopal Ramchurn, Stuart Russell, Marietje Schaake, Dawn Song, Alvaro Soto, Lee Tiedrich, Gaël Varoquaux, Andrew Yao, Ya-Qin Zhang

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.CY

Comments Available under the open government license at https://www.gov.uk/government/publications/international-scientific-report-on-the-safety-of-advanced-ai

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01252 2024-09-04 cs.CL cs.AI stat.ML 82%

Towards Scalable Automated Alignment of LLMs: A Survey

Boxi Cao, Keming Lu, Xinyu Lu, Jiawei Chen, Mengjie Ren, Hao Xiang, Peilin Liu, Yaojie Lu, Ben He, Xianpei Han, Le Sun, Hongyu Lin, Bowen Yu

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Paper List: https://github.com/cascip/awesome-auto-alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15645 2024-07-23 cs.CL cs.AI 82%

Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models

Joy He-Yueya, Wanjing Anya Ma, Kanishk Gandhi, Benjamin W. Domingue, Emma Brunskill, Noah D. Goodman

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Code and data: https://github.com/joyheyueya/psychometric-alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10635 2023-10-17 cs.LG cs.AI cs.CV 82%

Towards Scenario-based Safety Validation for Autonomous Trains with Deep Generative Models

Thomas Decker, Ananta R. Bhattarai, Michael Lebacher

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments International Conference on Computer Safety, Reliability, and Security 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.03057 2018-12-10 cs.CY cs.LG stat.ML 82%

Open Problems in Engineering and Quality Assurance of Safety Critical Machine Learning Systems

Hiroshi Kuwajima, Hirotoshi Yasuoka, Toshihiro Nakae

专题命中 其他安全 :safety(title,abstract);分类 cs.CY、cs.LG

Comments DISE1: Joint Workshop on Deep (or Machine) Learning for Safety-Critical Applications in Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07748 2026-05-11 cs.CL 81%

TextLDM: Language Modeling with Continuous Latent Diffusion

TextLDM:基于连续潜在扩散的语言建模

Jiaxiu Jiang, Jingjing Ren, Wenbo Li, Bo Wang, Haoze Sun, Yijun Yang, Jianhui Liu, Yanbing Zhang, Shenghe Zheng, Yuan Zhang, Haoyang Huang, Nan Duan, Wangmeng Zuo

机构 * Joy Future Academy(京东探索研究院) HIT(Harbin Institute of Technology) HKUST(GZ)(Hong Kong University of Science and Technology (Guangzhou))

专题命中 其他安全 :alignment(summary_cn,abstract);分类 cs.CL

AI总结 TextLDM将视觉潜在扩散框架应用于文本生成,通过Representation Alignment提升文本表示质量,在OpenWebText2上训练后优于现有扩散语言模型,匹配GPT-2性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15549 2026-05-01 cs.CL 81%

M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning Datasets

M-DaQ:用于指令微调数据集的多语言多样性与质量样本检索

Chunguang Zhao, Yilun Liu, Pufan Zeng, Yuanchang Luo, Shimin Tao, Minggui He, Weibin Meng, Song Xu, Chen Liu, Hongxia Ma, Li Zhang, Boxing Chen, Daimeng Wei

机构 * Huawei Technologies Ltd.(华为技术有限公司) University of Science and Technology of China(中国科学技术大学)

专题命中 其他安全 :alignment(summary_cn,abstract);分类 cs.CL

AI总结 M-DaQ通过联合优化指令-响应质量与跨语言语义多样性,构建高质量平衡训练数据,验证了多语言设置下的Superficial Alignment Hypothesis,并在18种语言上展示出超过60%的胜率。

Comments Accepted by SIGIR 2026 Short

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22104 2026-08-19 cs.AI cs.LG cs.LO cs.RO cs.SY eess.SY 版本更新 81%

Efficient Dynamic Shielding for Parametric Safety Specifications

面向参数化安全规范的高效动态防护机制

Davide Corsi, Kaushik Mallik, Andoni Rodriguez, Cesar Sanchez

机构 * University of California, Irvine(加州大学尔湾分校) IMDEA Software Institute(IMDEA软件研究所)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文针对参数化安全规范提出动态防护机制,其离线设计耗时数分钟,在线适配速度较暴力重新计算方法快最多5倍,可用于未知区域机器人导航以应对安全规范的动态变化。

Journal ref International Symposium on Automated Technology for Verification and Analysis (ATVA) 2025, pp. 157-179. Cham: Springer Nature Switzerland

详情

展开后加载摘要…

URL PDF HTML 收藏