arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2310.01708 2023-10-04 cs.CL 57%

Deciphering Diagnoses: How Large Language Models Explanations Influence Clinical Decision Making

D. Umerenkov, G. Zubkova, A. Nesterov

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00186 2023-10-03 cs.CV cs.AI 57%

Subject-driven Text-to-Image Generation via Apprenticeship Learning

Wenhu Chen, Hexiang Hu, Yandong Li, Nataniel Ruiz, Xuhui Jia, Ming-Wei Chang, William W. Cohen

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted at NeurIPS 2023. Model Service to be appear as Google Vertex AI - Instant Tuning (https://cloud.google.com/vertex-ai/docs/generative-ai/image/fine-tune-model). The link to demo video: https://www.youtube.com/watch?v=Q2xQ91D_dhM&t=2071s&ab_channel=GoogleCloud

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00374 2023-10-03 cs.CY 57%

Coordinated pausing: An evaluation-based coordination scheme for frontier AI developers

Jide Alaga, Jonas Schuett

专题命中 安全评测 :safety(abstract);分类 cs.CY

Comments 24 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15954 2023-09-29 cs.CV cs.LG 57%

The Devil is in the Details: A Deep Dive into the Rabbit Hole of Data Filtering

Haichao Yu, Yu Tian, Sateesh Kumar, Linjie Yang, Heng Wang

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 12 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15616 2023-09-28 cs.RO cs.AI 57%

Perception for Humanoid Robots

Arindam Roychoudhury, Shahram Khorshidi, Subham Agrawal, Maren Bennewitz

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 20 pages, 4 figures. To be published in Current Robotics Reports (Springer Nature)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11791 2023-09-28 cs.DL cs.CL 57%

SLHCat: Mapping Wikipedia Categories and Lists to DBpedia by Leveraging Semantic, Lexical, and Hierarchical Features

Zhaoyi Wang, Zhenyang Zhang, Jiaxin Qin, Mizuho Iwaihara

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments ICADL23

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.11671 2023-09-26 physics.med-ph cond-mat.mtrl-sci cs.LG physics.bio-ph physics.ins-det 57%

Capture Agent Free Biosensing using Porous Silicon Arrays and Machine Learning

Simon J. Ward, Tengfei Cao, Xiang Zhou, Catie Chang, Sharon M. Weiss

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 15 pages, 3 figures, 2 tables

Journal ref Biosensors 13 (2023) 1-12

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15324 2023-09-26 cs.AI 57%

Model evaluation for extreme risks

Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, Paul Christiano, Allan Dafoe

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Fixed typos; added citation

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13023 2023-09-26 cs.AI cs.CV 57%

Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images

Zeyu Lu, Di Huang, Lei Bai, Jingjing Qu, Chengyue Wu, Xihui Liu, Wanli Ouyang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12941 2023-09-25 cs.SE cs.AI 57%

Trusta: Reasoning about Assurance Cases with Formal Methods and Large Language Models

Zezhong Chen, Yuxin Deng, Wenjie Du

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 38 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12294 2023-09-22 cs.CL 57%

Reranking for Natural Language Generation from Logical Forms: A Study based on Large Language Models

Levon Haroutunian, Zhuang Li, Lucian Galescu, Philip Cohen, Raj Tumuluri, Gholamreza Haffari

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments IJCNLP-AACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11710 2023-09-22 cs.CL cs.CV 57%

ContextRef: Evaluating Referenceless Metrics For Image Description Generation

Elisa Kreiss, Eric Zelikman, Christopher Potts, Nick Haber

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09338 2023-09-19 cs.CL 57%

Performance of the Pre-Trained Large Language Model GPT-4 on Automated Short Answer Grading

Gerd Kortemeyer

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06062 2023-09-15 cs.LG cs.CV physics.geo-ph 57%

Selection of contributing factors for predicting landslide susceptibility using machine learning and deep learning models

Cheng Chen, Lei Fan

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Stochastic Environmental Research and Risk Assessment

Journal ref Stochastic Environmental Research and Risk Assessment, 13 September 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.10617 2023-09-11 cs.LG cs.SY eess.SY 57%

GPU-Accelerated Verification of Machine Learning Models for Power Systems

Samuel Chevalier, Ilgiz Murzakhanov, Spyros Chatzivasileiadis

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13387 2023-09-06 cs.CL 57%

Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, Timothy Baldwin

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments 18 pages, 9 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00172 2023-09-04 cs.AI 57%

Detecting Evidence of Organization in groups by Trajectories

T. F. Silva, J. E. B. Maia

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 17 pages, 16 figures, 3 algorithms, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03454 2023-08-30 cs.SE cs.CV cs.LG 57%

Benchmarking Robustness of AI-Enabled Multi-sensor Fusion Systems: Challenges and Opportunities

Xinyu Gao, Zhijie Wang, Yang Feng, Lei Ma, Zhenyu Chen, Baowen Xu

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments To appear in ESEC/FSE 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13228 2023-08-28 cs.HC cs.CY 57%

Meaningful XAI Based on User-Centric Design Methodology

Winston Maxwell, Bruno Dumas

专题命中 安全评测 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13116 2023-08-28 cs.CL 57%

Sentence Embedding Models for Ancient Greek Using Multilingual Knowledge Distillation

Kevin Krahn, Derrick Tate, Andrew C. Lamicela

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Paper accepted for publication at the First Workshop on Ancient Language Processing (ALP) 2023; 10 pages, 3 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03220 2023-08-28 cs.RO cs.AI 57%

Risk-Aware Reward Shaping of Reinforcement Learning Agents for Autonomous Driving

Lin-Chi Wu, Zengjie Zhang, Sofie Haesaert, Zhiqiang Ma, Zhiyong Sun

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03928 2023-08-23 cs.LG stat.AP 57%

Interpretable machine learning-accelerated seed treatment by nanomaterials for environmental stress alleviation

Hengjie Yu, Dan Luo, Sam F. Y. Li, Maozhen Qu, Da Liu, Yingchao He, Fang Cheng

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 30 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10441 2023-08-22 cs.AI cs.CV 57%

X-VoE: Measuring eXplanatory Violation of Expectation in Physical Events

Bo Dai, Linge Wang, Baoxiong Jia, Zeyu Zhang, Song-Chun Zhu, Chi Zhang, Yixin Zhu

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 19 pages, 16 figures, selected for an Oral presentation at ICCV 2023. Project link: https://pku.ai/publication/intuitive2023iccv/

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02955 2023-08-22 cs.SE cs.LG 57%

An Empirical Study of AI-based Smart Contract Creation

Rabimba Karanjai, Edward Li, Lei Xu, Weidong Shi

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Updated to address issues

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08655 2023-08-21 cs.LG 57%

Physics Informed Recurrent Neural Networks for Seismic Response Evaluation of Nonlinear Systems

Faisal Nissar Malik, James Ricles, Masoud Yari, Malik Arsala Nissar

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.03224 2023-08-16 cs.LG eess.SP 57%

Undersampling and Cumulative Class Re-decision Methods to Improve Detection of Agitation in People with Dementia

Zhidong Meng, Andrea Iaboni, Bing Ye, Kristine Newman, Alex Mihailidis, Zhihong Deng, Shehroz S. Khan

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06248 2023-08-14 cs.CV cs.LG 57%

FunnyBirds: A Synthetic Vision Dataset for a Part-Based Analysis of Explainable AI Methods

Robin Hesse, Simone Schaub-Meyer, Stefan Roth

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Accepted at ICCV 2023. Code: https://github.com/visinf/funnybirds

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.05120 2023-08-11 cs.LG 57%

Dynamic Model Agnostic Reliability Evaluation of Machine-Learning Methods Integrated in Instrumentation & Control Systems

Edward Chen, Han Bao, Nam Dinh

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments This paper was originally presented at the 13th Nuclear Plant Instrumentation, Control & Human-Machine Interface Technologies conference and was awarded best student paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04704 2023-08-11 cs.CR cs.LG 57%

A Feature Set of Small Size for the PDF Malware Detection

Ran Liu, Charles Nicholas

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Accepted for publication at the ACM SIGKDD & Annual KDD Conference workshop on Knowledge-infused Machine Learning, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.02934 2023-08-10 cs.LG 57%

Yggdrasil Decision Forests: A Fast and Extensible Decision Forests Library

Mathieu Guillame-Bert, Sebastian Bruch, Richard Stotz, Jan Pfeifer

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏