arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2311.07575 2023-11-14 cs.CV cs.AI cs.CL cs.LG 67%

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Ziyi Lin, Chris Liu, Renrui Zhang, Peng Gao, Longtian Qiu, Han Xiao, Han Qiu, Chen Lin, Wenqi Shao, Keqin Chen, Jiaming Han, Siyuan Huang, Yichi Zhang, Xuming He, Hongsheng Li, Yu Qiao

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Work in progress. Code and demos are released at https://github.com/Alpha-VLLM/LLaMA2-Accessory

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01918 2023-11-06 cs.CL cs.AI cs.LG 67%

Large Language Models Illuminate a Progressive Pathway to Artificial Healthcare Assistant: A Review

Mingze Yuan, Peng Bao, Jiajia Yuan, Yunhao Shen, Zifan Chen, Yi Xie, Jie Zhao, Yang Chen, Li Zhang, Lin Shen, Bin Dong

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 24 pages, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13064 2023-09-27 q-fin.GN cs.AI cs.CL cs.LG 67%

InvestLM: A Large Language Model for Investment using Financial Domain Instruction Tuning

Yi Yang, Yixuan Tang, Kar Yan Tam

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Link: https://github.com/AbaciNLP/InvestLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13442 2023-09-26 cs.LG cs.AI cs.CY 67%

How Do Drivers Behave at Roundabouts in a Mixed Traffic? A Case Study Using Machine Learning

Farah Abu Hamad, Rama Hasiba, Deema Shahwan, Huthaifa I. Ashqar

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.13243 2023-08-22 cs.LG cs.AI cs.CL cs.CV 67%

Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks

Tilman Räuker, Anson Ho, Stephen Casper, Dylan Hadfield-Menell

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02520 2023-07-18 cs.CL cs.AI cs.LG 67%

A Study of Situational Reasoning for Traffic Understanding

Jiarui Zhang, Filip Ilievski, Kaixin Ma, Aravinda Kollaa, Jonathan Francis, Alessandro Oltramari

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 11 pages, 6 figures, 5 tables, camera ready version of SIGKDD 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13000 2023-06-23 cs.CY cs.AI cs.CL 67%

Apolitical Intelligence? Auditing Delphi's responses on controversial political issues in the US

Jonathan H. Rystrøm

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08161 2023-06-19 cs.CL cs.AI cs.HC cs.IR cs.LG 67%

h2oGPT: Democratizing Large Language Models

Arno Candel, Jon McKinney, Philipp Singer, Pascal Pfeiffer, Maximilian Jeblick, Prithvi Prabhu, Jeff Gambera, Mark Landry, Shivam Bansal, Ryan Chesler, Chun Ming Lee, Marcos V. Conde, Pasha Stetsenko, Olivier Grellier, SriSatish Ambati

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Work in progress by H2O.ai, Inc

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02622 2023-06-06 cs.LG cs.AI cs.CL 67%

What Makes Entities Similar? A Similarity Flooding Perspective for Multi-sourced Knowledge Graph Embeddings

Zequn Sun, Jiacheng Huang, Xiaozhou Xu, Qijin Chen, Weijun Ren, Wei Hu

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted in the 40th International Conference on Machine Learning (ICML 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08726 2023-06-01 cs.CL cs.AI cs.IR cs.LG 67%

RARR: Researching and Revising What Language Models Say, Using Language Models

Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y. Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, Kelvin Guu

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16504 2023-05-29 cs.CL cs.AI cs.LG 67%

On the Tool Manipulation Capability of Open-source Large Language Models

Qiantong Xu, Fenglu Hong, Bo Li, Changran Hu, Zhengyu Chen, Jian Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14095 2023-04-28 cs.NI 67%

Securing Autonomous Air Traffic Management: Blockchain Networks Driven by Explainable AI

Louise Axon, Dimitrios Panagiotakopoulos, Samuel Ayo, Carolina Sanchez-Hernandez, Yan Zong, Simon Brown, Lei Zhang, Michael Goldsmith, Sadie Creese, Weisi Guo

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

Comments under review in IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.03052 2023-01-10 cs.LG cs.AI cs.CY 67%

AI Maintenance: A Robustness Perspective

Pin-Yu Chen, Payel Das

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted to IEEE Computer Magazine. To be published in 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00355 2023-01-06 cs.CL cs.AI cs.CY 67%

Second Thoughts are Best: Learning to Re-Align With Human Values from Text Edits

Ruibo Liu, Chenyan Jia, Ge Zhang, Ziyu Zhuang, Tony X Liu, Soroush Vosoughi

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments In proceedings of the 36th Conference on Neural Information Processing Systems (NeurIPS 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00544 2022-12-02 cs.RO 67%

Towards Explainability in Modular Autonomous Vehicle Software

Hongrui Zheng, Zirui Zang, Shuo Yang, Rahul Mangharam

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12757 2022-11-24 cs.LG cs.AI cs.CY 67%

FAIRification of MLC data

Ana Kostovska, Jasmin Bogatinovski, Andrej Treven, Sašo Džeroski, Dragi Kocev, Panče Panov

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments This paper was accepted ECML PKDD 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10693 2022-10-20 cs.CL cs.AI cs.LG 67%

Robustness of Demonstration-based Learning Under Limited Data Scenario

Hongxin Zhang, Yanzhe Zhang, Ruiyi Zhang, Diyi Yang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 14 pages, EMNLP 2022 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04995 2022-10-12 cs.LG cs.AI cs.CY 67%

FEAMOE: Fair, Explainable and Adaptive Mixture of Experts

Shubham Sharma, Jette Henderson, Joydeep Ghosh

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.05862 2022-09-21 cs.CY cs.AI cs.LG 67%

X-Risk Analysis for AI Research

Dan Hendrycks, Mantas Mazeika

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.09079 2022-08-22 cs.LG cs.AI cs.CV cs.CY 67%

A Multi-Modal Wildfire Prediction and Personalized Early-Warning System Based on a Novel Machine Learning Framework

Rohan Tan Bhowmik

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08080 2022-08-18 cs.AI cs.CL cs.CV cs.LG cs.MM 67%

Multimodal Lecture Presentations Dataset: Understanding Multimodality in Educational Slides

Dong Won Lee, Chaitanya Ahuja, Paul Pu Liang, Sanika Natu, Louis-Philippe Morency

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.05811 2022-07-14 cs.LG cs.AI cs.CY 67%

Revealing Unfair Models by Mining Interpretable Evidence

Mohit Bajaj, Lingyang Chu, Vittorio Romaniello, Gursimran Singh, Jian Pei, Zirui Zhou, Lanjun Wang, Yong Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.12725 2022-06-28 cs.CV 67%

Empirical Evaluation of Physical Adversarial Patch Attacks Against Overhead Object Detection Models

Gavin S. Hartnett, Li Ang Zhang, Caolionn O'Connell, Andrew J. Lohn, Jair Aguirre

专题命中 安全评测 :safety(abstract);AI safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06490 2022-03-22 cs.CL cs.AI cs.LG 67%

Dict-BERT: Enhancing Language Model Pre-training with Dictionary

Wenhao Yu, Chenguang Zhu, Yuwei Fang, Donghan Yu, Shuohang Wang, Yichong Xu, Michael Zeng, Meng Jiang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2022 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.02776 2022-02-08 cs.AI cs.CY cs.HC cs.LG 67%

Human rights, democracy, and the rule of law assurance framework for AI systems: A proposal

David Leslie, Christopher Burr, Mhairi Aitken, Michael Katell, Morgan Briggs, Cami Rincon

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 341 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.07754 2021-11-09 cs.AI cs.CY cs.LG stat.ML 67%

Counterfactual Explanations as Interventions in Latent Space

Riccardo Crupi, Alessandro Castelnovo, Daniele Regoli, Beatriz San Miguel Gonzalez

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 34 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.11036 2021-06-22 cs.CY cs.AI cs.LG 67%

Know Your Model (KYM): Increasing Trust in AI and Machine Learning

Mary Roszel, Robert Norvill, Jean Hilger, Radu State

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.09539 2020-12-18 cs.LO 67%

Online Shielding for Stochastic Systems

Bettina Könighofer, Julian Rudolf, Alexander Palmisano, Martin Tappler, Roderick Bloem

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

Comments 18 Pages, 6 Figures, under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.01318 2019-04-03 cs.CV 67%

Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents

Christian Rupprecht, Cyril Ibrahim, Christopher J. Pal

专题命中 安全评测 :safety(abstract);AI safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.08757 2018-10-09 cs.CV 67%

Siamese Generative Adversarial Privatizer for Biometric Data

Witold Oleszkiewicz, Peter Kairouz, Karol Piczak, Ram Rajagopal, Tomasz Trzcinski

专题命中 安全评测 :safety(abstract);AI safety(abstract)

Comments Paper accepted to ACCV 2018 (Asian Conference on Computer Vision)

详情

展开后加载摘要…

URL PDF HTML 收藏