arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2112.00552 2022-06-20 cs.LG cs.AI cs.LO 62%

SaDe: Learning Models that Provably Satisfy Domain Constraints

Kshitij Goyal, Sebastijan Dumancic, Hendrik Blockeel

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.15509 2022-06-01 cs.CV cs.AI cs.CL 62%

ADAPT: Vision-Language Navigation with Modality-Aligned Action Prompts

Bingqian Lin, Yi Zhu, Zicong Chen, Xiwen Liang, Jianzhuang Liu, Xiaodan Liang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.13856 2022-05-30 cs.CL cs.AI 62%

Lightweight Cross-Lingual Sentence Representation Learning

Zhuoyuan Mao, Prakhar Gupta, Pei Wang, Chenhui Chu, Martin Jaggi, Sadao Kurohashi

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ACL 2021 main conference; modified Eq. (2)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10123 2022-05-23 cs.AI cs.LG 62%

Lifelong Personal Context Recognition

Andrea Bontempelli, Marcelo Rodas Britez, Xiaoyue Li, Haonan Zhao, Luca Erculiani, Stefano Teso, Andrea Passerini, Fausto Giunchiglia

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.00948 2022-05-23 cs.CL cs.LG 62%

Unsupervised Out-of-Domain Detection via Pre-trained Transformers

Keyang Xu, Tongzheng Ren, Shikun Zhang, Yihao Feng, Caiming Xiong

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

Comments Accepted by ACL 2021. Code is available at https://github.com/rivercold/BERT-unsupervised-OOD

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.07623 2022-05-17 cs.AI cs.LG 62%

Model Agnostic Local Explanations of Reject

André Artelt, Roel Visser, Barbara Hammer

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: text overlap with arXiv:2202.07244

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.07233 2022-05-17 cs.CL cs.AI 62%

Mitigating Toxic Degeneration with Empathetic Data: Exploring the Relationship Between Toxicity and Empathy

Allison Lahnala, Charles Welch, Béla Neuendorf, Lucie Flek

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.14782 2022-05-05 cs.CL cs.LG 62%

When is BERT Multilingual? Isolating Crucial Ingredients for Cross-lingual Transfer

Ameet Deshpande, Partha Talukdar, Karthik Narasimhan

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted at NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01057 2022-05-03 cs.LG cs.AI 62%

Causal Discovery on the Effect of Antipsychotic Drugs on Delirium Patients in the ICU using Large EHR Dataset

Riddhiman Adib, Md Osman Gani, Sheikh Iqbal Ahamed, Mohammad Adibuzzaman

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00210 2022-05-03 cs.SE cs.AI cs.LG 62%

Software Testing for Machine Learning

Dusica Marijan, Arnaud Gotlieb

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 34(09), 13576-13582 (2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12716 2022-04-28 cs.CL cs.AI 62%

UBERT: A Novel Language Model for Synonymy Prediction at Scale in the UMLS Metathesaurus

Thilini Wijesiriwardene, Vinh Nguyen, Goonmeet Bajaj, Hong Yung Yip, Vishesh Javangula, Yuqing Mao, Kin Wah Fung, Srinivasan Parthasarathy, Amit P. Sheth, Olivier Bodenreider

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05210 2022-04-12 cs.CL cs.AI 62%

Bridging the Gap between Language Models and Cross-Lingual Sequence Labeling

Nuo Chen, Linjun Shou, Ming Gong, Jian Pei, Daxin Jiang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05091 2022-04-12 cs.AI cs.CL 62%

Linguistic communication as (inverse) reward design

Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Thomas L. Griffiths, Dylan Hadfield-Menell

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 6 pages, 3 figures. Accepted at Learning from Natural Language Supervision workshop (ACL 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.06717 2022-04-08 cs.LG cs.AI stat.ML 62%

MoËT: Mixture of Expert Trees and its Application to Verifiable Reinforcement Learning

Marko Vasic, Andrija Petrovic, Kaiyuan Wang, Mladen Nikolic, Rishabh Singh, Sarfraz Khurshid

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref Neural Networks, Volume 151, 2022, Pages 34-47

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.01385 2022-04-06 cs.CL cs.LG 62%

Aligned Weight Regularizers for Pruning Pretrained Neural Networks

James O' Neill, Sourav Dutta, Haytham Assem

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted to ACL Findings 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.01984 2022-04-05 cs.SE cs.AI cs.LG 62%

Software Engineering for AI-Based Systems: A Survey

Silverio Martínez-Fernández, Justus Bogner, Xavier Franch, Marc Oriol, Julien Siebert, Adam Trendowicz, Anna Maria Vollmer, Stefan Wagner

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted in ACM Transactions on Software Engineering and Methodology (TOSEM). For its published version refer to the Journal of ACM TOSEM

Journal ref ACM Trans. Softw. Eng. Methodol. 31, 2, Article 37e (March 2022), 59 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06386 2022-03-31 cs.CL cs.AI cs.CV 62%

Enabling Multimodal Generation on CLIP via Vision-Language Knowledge Distillation

Wenliang Dai, Lu Hou, Lifeng Shang, Xin Jiang, Qun Liu, Pascale Fung

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.07785 2022-03-16 cs.CL cs.AI 62%

The Ghost in the Machine has an American accent: value conflict in GPT-3

Rebecca L Johnson, Giada Pistilli, Natalia Menédez-González, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, Donald Jay Bertulfo

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments There are a total of 15 pages of the PDF including 8 pages of the main manuscript, 3 pages of references, and 4 pages of appendices. The paper is currently under review by a conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.04563 2022-03-10 cs.RO cs.AI cs.CV cs.LG 62%

MLNav: Learning to Safely Navigate on Martian Terrains

Shreyansh Daftry, Neil Abcouwer, Tyler Del Sesto, Siddarth Venkatraman, Jialin Song, Lucas Igel, Amos Byon, Ugo Rosolia, Yisong Yue, Masahiro Ono

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments IEEE Robotics and Automation Letters (RA-L) and ICRA 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02959 2022-03-08 cs.RO cs.AI cs.CV cs.HC cs.LG 62%

A Perspective on Robotic Telepresence and Teleoperation using Cognition: Are we there yet?

Hrishav Bakul Barua, Ashis Sau, Ruddra dev Roychoudhury

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.12453 2022-02-17 cs.AI cs.HC cs.LG cs.NE cs.RO 62%

Fanoos: Multi-Resolution, Multi-Strength, Interactive Explanations for Learned Systems

David Bayani, Stefan Mitsch

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 60 pages, 20 pages main body, 128 references, 3 figures, 5 tables, 12 pseudocode blocks Update 24 Sep. 2020 : Added a pointer to further, external content: Append Section E. Update 20 Mar. 2021: Substantial additions. Further explanations of process, with far more pseudocode. Some corrections to a previous description; see errata section. Also briefly describe a few implemented extensions

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.03544 2022-02-15 cs.LG cs.AI stat.ML 62%

The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Alexander Pan, Kush Bhatia, Jacob Steinhardt

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments ICLR 2022; 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.11222 2022-02-08 cs.LG cs.AI 62%

Is High Variance Unavoidable in RL? A Case Study in Continuous Control

Johan Bjorck, Carla P. Gomes, Kilian Q. Weinberger

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to ICLR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.01374 2022-02-04 cs.CL cs.LG 62%

mSLAM: Massively multilingual joint pre-training for speech and text

Ankur Bapna, Colin Cherry, Yu Zhang, Ye Jia, Melvin Johnson, Yong Cheng, Simran Khanuja, Jason Riesa, Alexis Conneau

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.10349 2022-01-26 cs.CR cs.AI cs.LG 62%

Roadmap for Cybersecurity in Autonomous Vehicles

Vipin Kumar Kukkala, Sooryaa Vignesh Thiruloga, Sudeep Pasricha

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.05877 2022-01-19 cs.LG cs.AI 62%

A Framework for Pedestrian Sub-classification and Arrival Time Prediction at Signalized Intersection Using Preprocessed Lidar Data

Tengfeng Lin, Zhixiong Jin, Seongjin Choi, Hwasoo Yeo

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 15 pages, 11 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10329 2021-10-22 cs.CL cs.LG 62%

SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training

Ankur Bapna, Yu-an Chung, Nan Wu, Anmol Gulati, Ye Jia, Jonathan H. Clark, Melvin Johnson, Jason Riesa, Alexis Conneau, Yu Zhang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09246 2021-10-19 cs.LG cs.AI 62%

Single Layer Predictive Normalized Maximum Likelihood for Out-of-Distribution Detection

Koby Bibas, Meir Feder, Tal Hassner

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.11914 2021-10-15 cs.LG cs.AI cs.CV cs.SC 62%

EXplainable Neural-Symbolic Learning (X-NeSyL) methodology to fuse deep learning representations with expert knowledge graphs: the MonuMAI cultural heritage use case

Natalia Díaz-Rodríguez, Alberto Lamas, Jules Sanchez, Gianni Franchi, Ivan Donadello, Siham Tabik, David Filliat, Policarpo Cruz, Rosana Montes, Francisco Herrera

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.13222 2021-09-28 cs.CL cs.LG 62%

Using Pause Information for More Accurate Entity Recognition

Sahas Dendukuri, Pooja Chitkara, Joel Ruben Antony Moniz, Xiao Yang, Manos Tsagkias, Stephen Pulman

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted at the 3rd Workshop on NLP for Conversational AI (EMNLP 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏