arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8064 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8064 篇

2206.08823 2023-11-01 cs.CL 57%

Language with Vision: a Study on Grounded Word and Sentence Embeddings

Hassan Shahmohammadi, Maria Heitmeier, Elnaz Shafaei-Bajestan, Hendrik P. A. Lensch, Harald Baayen

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19449 2023-10-31 cs.AI 57%

Large-Scale Application of Fault Injection into PyTorch Models -- an Extension to PyTorchFI for Validation Efficiency

Ralf Graafe, Qutub Syed Sha, Florian Geissler, Michael Paulitsch

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments accepted in DSN2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.17133 2023-10-31 cs.CL cs.CV 57%

Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question Answering

Weizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca, Bill Byrne

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments To appear at NeurIPS 2023. This is the camera-ready version. We fixed some numbers and added more experiments to address reviewers' comments

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15640 2023-10-31 cs.LG cs.CV 57%

Characterizing Out-of-Distribution Error via Optimal Transport

Yuzhe Lu, Yilong Qin, Runtian Zhai, Andrew Shen, Ketong Chen, Zhenlin Wang, Soheil Kolouri, Simon Stepputtis, Joseph Campbell, Katia Sycara

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15109 2023-10-24 cs.CL 57%

GRENADE: Graph-Centric Language Model for Self-Supervised Representation Learning on Text-Attributed Graphs

Yichuan Li, Kaize Ding, Kyumin Lee

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Findings of EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12074 2023-10-24 cs.CL 57%

Towards Safer Operations: An Expert-involved Dataset of High-Pressure Gas Incidents for Preventing Future Failures

Shumpei Inoue, Minh-Tien Nguyen, Hiroki Mizokuchi, Tuan-Anh D. Nguyen, Huu-Hiep Nguyen, Dung Tien Le

专题命中 其他安全 :safety(abstract);分类 cs.CL

Comments Accepted by EMNLP 2023 (The Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14485 2023-10-24 cs.RO cs.AI cs.SY eess.SY 57%

Intelligent Escape of Robotic Systems: A Survey of Methodologies, Applications, and Challenges

Junfei Li, Simon X. Yang

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments This paper is accepted by Journal of Intelligent and Robotic Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14092 2023-10-24 cs.RO cs.AI 57%

Learning Reward for Physical Skills using Large Language Model

Yuwei Zeng, Yiqing Xu

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments CoRL 2023, LangRob workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13303 2023-10-23 cs.CL 57%

Towards Unsupervised Recognition of Token-level Semantic Differences in Related Documents

Jannis Vamvas, Rico Sennrich

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12994 2023-10-23 q-bio.NC cs.AI 57%

Dimensions of Disagreement: Unpacking Divergence and Misalignment in Cognitive Science and Artificial Intelligence

Kerem Oktar, Ilia Sucholutsky, Tania Lombrozo, Thomas L. Griffiths

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10661 2023-10-18 cs.CR cs.LG 57%

TII-SSRC-23 Dataset: Typological Exploration of Diverse Traffic Patterns for Intrusion Detection

Dania Herzalla, Willian T. Lunardi, Martin Andreoni Lopez

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10262 2023-10-17 cs.CL 57%

Enhancing Interpretability using Human Similarity Judgements to Prune Word Embeddings

Natalia Flechas Manrique, Wanqian Bao, Aurelie Herbelot, Uri Hasson

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted for presentation at the BlackboxNLP workshop at EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08034 2023-10-13 cs.HC cs.AI cs.RO 57%

Receive, Reason, and React: Drive as You Say with Large Language Models in Autonomous Vehicles

Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, Ziran Wang

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:2309.10228

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08021 2023-10-11 cs.CV cs.AI 57%

Vision-based Analysis of Driver Activity and Driving Performance Under the Influence of Alcohol

Ross Greer, Akshay Gopalkrishnan, Sumega Mandadi, Pujitha Gunaratne, Mohan M. Trivedi, Thomas D. Marcotte

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments Withdrawn at the request of industry research collaborators, per contract agreement

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13160 2023-10-11 cs.CL 57%

Can ChatGPT Defend its Belief in Truth? Evaluating LLM Reasoning via Debate

Boshi Wang, Xiang Yue, Huan Sun

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP-23 (findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05878 2023-10-10 cs.LG 57%

A Machine Learning Approach to Predicting Single Event Upsets

Archit Gupta, Chong Yock Eng, Deon Lim Meng Wee, Rashna Analia Ahmed, See Min Sim

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05824 2023-10-10 cs.CL 57%

Terminology-Aware Translation with Constrained Decoding and Large Language Model Prompting

Nikolay Bogoychev, Pinzhen Chen

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments WMT 2023 Terminology Translation Task

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04726 2023-10-10 cs.CL 57%

Zero-shot Cross-lingual Transfer without Parallel Corpus

Yuyang Zhang, Xiaofeng Han, Baojun Wang

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04605 2023-10-10 cs.LG math.OC 57%

Learning Optimal Power Flow Value Functions with Input-Convex Neural Networks

Andrew Rosemberg, Mathieu Tanneau, Bruno Fanzeres, Joaquim Garcia, Pascal Van Hentenryck

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06235 2023-10-10 cs.LG q-bio.BM 57%

Multimodal Molecular Pretraining via Modality Blending

Qiying Yu, Yudi Zhang, Yuyan Ni, Shikun Feng, Yanyan Lan, Hao Zhou, Jingjing Liu

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04288 2023-10-09 eess.SY cs.AI cs.FL cs.SY 57%

Searching for Optimal Runtime Assurance via Reachability and Reinforcement Learning

Kristina Miller, Christopher K. Zeitler, William Shen, Kerianne Hobbs, Sayan Mitra, John Schierman, Mahesh Viswanathan

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03999 2023-10-09 cs.LG cs.SE 57%

Runtime Monitoring DNN-Based Perception

Chih-Hong Cheng, Michael Luttenberger, Rongjie Yan

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03269 2023-10-06 q-bio.BM cs.CL 57%

InstructProtein: Aligning Human and Protein Language via Knowledge Instruction

Zeyuan Wang, Qiang Zhang, Keyan Ding, Ming Qin, Xiang Zhuang, Xiaotong Li, Huajun Chen

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03217 2023-10-06 cs.LG 57%

Formal and Practical Elements for the Certification of Machine Learning Systems

Jean-Guillaume Durand, Arthur Dubois, Robert J. Moss

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments Best of Conference at the 2023 Digital Avionics Systems Conference (DASC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00653 2023-10-03 cs.CV cs.AI 57%

Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Tianyu Yu, Jinyi Hu, Yuan Yao, Haoye Zhang, Yue Zhao, Chongyi Wang, Shan Wang, Yinxv Pan, Jiao Xue, Dahai Li, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14292 2023-09-26 cs.AR cs.ET cs.LG 57%

On the Non-Associativity of Analog Computations

Lisa Kuhn, Bernhard Klein, Holger Fröning

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments Published at the ECML PKDD Conference 2023, at the 4th Workshop on IoT, Edge, and Mobile for Embedded Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09987 2023-09-20 cs.LG cs.CV 57%

TCGF: A unified tensorized consensus graph framework for multi-view representation learning

Xiangzhu Meng, Wei Wei, Qiang Liu, Shu Wu, Liang Wang

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.07742 2023-09-15 cs.LG cs.HC 57%

Interpretability is in the Mind of the Beholder: A Causal Framework for Human-interpretable Representation Learning

Emanuele Marconato, Andrea Passerini, Stefano Teso

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04616 2023-09-13 cs.LG cs.SE 57%

Knowledge Distillation-Empowered Digital Twin for Anomaly Detection

Qinghua Xu, Shaukat Ali, Tao Yue, Zaimovic Nedim, Inderjeet Singh

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05608 2023-09-12 cs.CL cs.CE 57%

Incorporating Pre-trained Model Prompting in Multimodal Stock Volume Movement Prediction

Ruibo Chen, Zhiyuan Zhang, Yi Liu, Ruihan Bao, Keiko Harimoto, Xu Sun

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 9 pages, 3 figures, 7 tables. Accepted by 2023 KDD Workshop on Machine Learning in Finance

详情

展开后加载摘要…

URL PDF HTML 收藏