arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9400 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9400 篇

2403.13313 2024-03-21 cs.AI cs.CL 81%

Polaris: A Safety-focused LLM Constellation Architecture for Healthcare

Subhabrata Mukherjee, Paul Gamble, Markel Sanz Ausin, Neel Kant, Kriti Aggarwal, Neha Manjunath, Debajyoti Datta, Zhengliang Liu, Jiayuan Ding, Sophia Busacca, Cezanne Bianco, Swapnil Sharma, Rae Lasko, Michelle Voisard, Sanchay Harneja, Darya Filippova, Gerry Meixiong, Kevin Cha, Amir Youssefi, Meyhaa Buvanesh, Howard Weingram, Sebastian Bierman-Lytle, Harpreet Singh Mangat, Kim Parikh, Saad Godil, Alex Miller

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09015 2024-02-26 cs.CL cs.AI 81%

Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications

Negar Arabzadeh, Julia Kiseleva, Qingyun Wu, Chi Wang, Ahmed Awadallah, Victor Dibia, Adam Fourney, Charles Clarke

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13635 2024-02-22 cs.LG cs.AI 81%

The METRIC-framework for assessing data quality for trustworthy AI in medicine: a systematic review

Daniel Schwabe, Katinka Becker, Martin Seyferth, Andreas Klaß, Tobias Schäffter

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12417 2024-02-21 cs.LG cs.AI 81%

Predicting trucking accidents with truck drivers 'safety climate perception across companies: A transfer learning approach

Kailai Sun, Tianxiang Lan, Say Hong Kam, Yang Miang Goh, Yueng-Hsiang Huang

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments submitted to journal: accident analysis and prevention

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02055 2024-02-06 cs.LG cs.AI 81%

Variance Alignment Score: A Simple But Tough-to-Beat Data Selection Method for Multimodal Contrastive Learning

Yiping Wang, Yifang Chen, Wendan Yan, Kevin Jamieson, Simon Shaolei Du

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 17 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01096 2024-02-05 cs.LG cs.AI cs.CR cs.DC 81%

Trustworthy Distributed AI Systems: Robustness, Privacy, and Governance

Wenqi Wei, Ling Liu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments Manuscript accepted to ACM Computing Surveys

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.18058 2024-02-01 cs.CL cs.LG 81%

LongAlign: A Recipe for Long Context Alignment of Large Language Models

Yushi Bai, Xin Lv, Jiajie Zhang, Yuze He, Ji Qi, Lei Hou, Jie Tang, Yuxiao Dong, Juanzi Li

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12474 2024-01-24 cs.CL cs.LG 81%

Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment

Keming Lu, Bowen Yu, Chang Zhou, Jingren Zhou

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02779 2024-01-18 eess.IV cs.AI cs.CV cs.LG 81%

A Dempster-Shafer approach to trustworthy AI with application to fetal brain MRI segmentation

Lucas Fidon, Michael Aertsen, Florian Kofler, Andrea Bink, Anna L. David, Thomas Deprest, Doaa Emam, Frédéric Guffens, András Jakab, Gregor Kasprian, Patric Kienast, Andrew Melbourne, Bjoern Menze, Nada Mufti, Ivana Pogledic, Daniela Prayer, Marlene Stuempflen, Esther Van Elslander, Sébastien Ourselin, Jan Deprest, Tom Vercauteren

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments Published in IEEE TPAMI. Minor revision compared to the previous version

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16101 2023-11-28 cs.CV cs.CL cs.LG 81%

How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Haoqin Tu, Chenhang Cui, Zijun Wang, Yiyang Zhou, Bingchen Zhao, Junlin Han, Wangchunshu Zhou, Huaxiu Yao, Cihang Xie

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.LG

Comments H.T., C.C., and Z.W. contribute equally. Work done during H.T. and Z.W.'s internship at UCSC, and C.C. and Y.Z.'s internship at UNC

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10819 2023-11-28 cs.CL cs.AI 81%

Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection

Zekun Li, Baolin Peng, Pengcheng He, Xifeng Yan

专题命中 安全评测 :prompt injection(title,abstract);分类 cs.CL、cs.AI

Comments The data and code can be found at https://github.com/Leezekun/instruction-following-robustness-eval

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02457 2023-11-16 cs.CL cs.CY 81%

The Empty Signifier Problem: Towards Clearer Paradigms for Operationalising "Alignment" in Large Language Models

Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, Scott A. Hale

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.CY

Comments Socially Responsible Language Modelling Research (SoLaR) @ NeurIPs 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03795 2023-11-08 cs.AI cs.HC cs.LG 81%

AI-Supported Assessment of Load Safety

Julius Schöning, Niklas Kruse

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 9 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12443 2023-10-20 cs.IR cs.AI cs.CL 81%

Know Where to Go: Make LLM a Relevant, Responsible, and Trustworthy Searcher

Xiang Shi, Jiawei Liu, Yinpeng Liu, Qikai Cheng, Wei Lu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments 14 pages, 4 figures, under peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08215 2023-10-13 cs.LG cs.AI 81%

Trustworthy Machine Learning

Bálint Mucsányi, Michael Kirchhof, Elisa Nguyen, Alexander Rubinstein, Seong Joon Oh

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 373 pages, textbook at the University of Tübingen

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12324 2023-09-25 cs.CY cs.LG cs.SY eess.SY 81%

Aviation Safety Risk Analysis and Flight Technology Assessment Issues

Shuanghe Liu

专题命中 安全评测 :safety(title,abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09450 2023-09-19 cs.CY cs.AI cs.HC 81%

Are You Worthy of My Trust?: A Socioethical Perspective on the Impacts of Trustworthy AI Systems on the Environment and Human Society

Jamell Dacon

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12315 2023-08-31 cs.LG cs.AI 81%

Trustworthy Representation Learning Across Domains

Ronghang Zhu, Dongliang Guo, Daiqing Qi, Zhixuan Chu, Xiang Yu, Sheng Li

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 38 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04445 2023-08-10 cs.LG cs.AI 81%

Getting from Generative AI to Trustworthy AI: What LLMs might learn from Cyc

Doug Lenat, Gary Marcus

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 21 pages, 1 Figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00076 2023-08-02 cs.AI cs.LG stat.ML 81%

Crowd Safety Manager: Towards Data-Driven Active Decision Support for Planning and Control of Crowd Events

Panchamy Krishnakumari, Sascha Hoogendoorn-Lanser, Jeroen Steenbakkers, Serge Hoogendoorn

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments Submitted to TRB Annual Meeting 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16851 2023-08-01 cs.LG cs.AI 81%

Towards Trustworthy and Aligned Machine Learning: A Data-centric Survey with Causality Perspectives

Haoyang Liu, Maheep Chaudhary, Haohan Wang

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 47 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04683 2023-07-11 cs.CL cs.AI 81%

CORE-GPT: Combining Open Access research and large language models for credible, trustworthy question answering

David Pride, Matteo Cancellieri, Petr Knoth

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments 12 pages, accepted submission to TPDL2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11507 2023-06-21 cs.CL cs.AI 81%

TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Yue Huang, Qihui Zhang, Philip S. Y, Lichao Sun

专题命中 安全评测 :trustworthy(title);alignment(abstract);分类 cs.CL、cs.AI

Comments We are currently expanding this work and welcome collaborators!

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08728 2023-06-16 cs.LG cs.AI eess.SP 81%

Towards trustworthy seizure onset detection using workflow notes

Khaled Saab, Siyi Tang, Mohamed Taha, Christopher Lee-Messer, Christopher Ré, Daniel Rubin

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07993 2023-06-16 cs.CR cs.AI cs.LG 81%

Trustworthy Artificial Intelligence Framework for Proactive Detection and Risk Explanation of Cyber Attacks in Smart Grid

Md. Shirajum Munir, Sachin Shetty, Danda B. Rawat

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments Submitted for peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18307 2023-05-31 cs.CY cs.AI 81%

Certification Labels for Trustworthy AI: Insights From an Empirical Mixed-Method Study

Nicolas Scharowski, Michaela Benk, Swen J. Kühne, Léane Wettstein, Florian Brühlmann

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14384 2023-05-25 cs.LG cs.AI cs.CR cs.CV 81%

Adversarial Nibbler: A Data-Centric Challenge for Improving the Safety of Text-to-Image Models

Alicia Parrish, Hannah Rose Kirk, Jessica Quaye, Charvi Rastogi, Max Bartolo, Oana Inel, Juan Ciro, Rafael Mosquera, Addison Howard, Will Cukierski, D. Sculley, Vijay Janapa Reddi, Lora Aroyo

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00470 2023-05-04 cs.LG cs.CY cs.GT 81%

Reward Systems for Trustworthy Medical Federated Learning

Konstantin D. Pandl, Florian Leiser, Scott Thiebes, Ali Sunyaev

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13518 2023-03-24 cs.CV cs.AI cs.LG 81%

Three ways to improve feature alignment for open vocabulary detection

Relja Arandjelović, Alex Andonian, Arthur Mensch, Olivier J. Hénaff, Jean-Baptiste Alayrac, Andrew Zisserman

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07778 2023-03-15 cs.LG cs.AI 81%

GANN: Graph Alignment Neural Network for Semi-Supervised Learning

Linxuan Song, Wenxuan Tu, Sihang Zhou, Xinwang Liu, En Zhu

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏