arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9400 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9400 篇

2411.02094 2024-11-05 cs.HC cs.AI cs.LG 81%

Alignment-Based Adversarial Training (ABAT) for Improving the Robustness and Accuracy of EEG-Based BCIs

Xiaoqing Chen, Ziwei Wang, Dongrui Wu

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Journal ref IEEE Trans. on Neural Systems and Rehabilitation Engineering, 32:1703-1714, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13911 2024-11-05 cs.CV cs.AI cs.CL 81%

TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment

Wei Li, Hehe Fan, Yongkang Wong, Mohan Kankanhalli, Yi Yang

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2024 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12649 2024-11-04 cs.LG cs.AI cs.CV stat.ML 81%

Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation Models

Hengyi Wang, Shiwei Tan, Hao Wang

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments Proceedings of the 41st International Conference on Machine Learning (ICML 2024)

Journal ref PMLR 235:51502-51522, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23703 2024-11-01 cs.LG cs.CL 81%

OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models

Junda Wu, Xintong Li, Ruoyu Wang, Yu Xia, Yuxin Xiong, Jianing Wang, Tong Yu, Xiang Chen, Branislav Kveton, Lina Yao, Jingbo Shang, Julian McAuley

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21695 2024-10-30 cs.CL cs.LG 81%

CFSafety: Comprehensive Fine-grained Safety Assessment for LLMs

Zhihao Liu, Chenhui Hu

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21346 2024-10-30 cs.LG cs.AI 81%

Towards Trustworthy Machine Learning in Production: An Overview of the Robustness in MLOps Approach

Firas Bayram, Bestoun S. Ahmed

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11446 2024-10-28 cs.LG cs.AI stat.ML 81%

Exploration of the Rashomon Set Assists Trustworthy Explanations for Medical Data

Katarzyna Kobylińska, Mateusz Krzyziński, Rafał Machowicz, Mariusz Adamek, Przemysław Biecek

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16606 2024-10-23 cs.LG cs.AI 81%

GALA: Graph Diffusion-based Alignment with Jigsaw for Source-free Domain Adaptation

Junyu Luo, Yiyang Gu, Xiao Luo, Wei Ju, Zhiping Xiao, Yusheng Zhao, Jingyang Yuan, Ming Zhang

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10093 2024-10-15 cs.CL cs.LG 81%

How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective

Teng Xiao, Mingxiao Li, Yige Yuan, Huaisheng Zhu, Chao Cui, Vasant G Honavar

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments EMNLP 2024 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16700 2024-10-08 cs.CV cs.CL cs.LG 81%

Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs

Mustafa Shukor, Matthieu Cord

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments NeurIPS 2024. Code: https://github.com/mshukor/ima-lmms. Project page: https://ima-lmms.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01927 2024-10-04 cs.CY cs.AI econ.GN q-fin.EC 81%

Risk Alignment in Agentic AI Systems

Hayley Clatterbuck, Clinton Castro, Arvo Muñoz Morán

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00741 2024-10-01 cs.AI cs.LG cs.RO 81%

Diffusion Models for Offline Multi-agent Reinforcement Learning with Safety Constraints

Jianuo Huang

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments The experiment and method plan are abolished and need to be redesigned

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17819 2024-09-27 cs.CL cs.AI 81%

Inference-Time Language Model Alignment via Integrated Value Guidance

Zhixuan Liu, Zhanhui Zhou, Yuanfu Wang, Chao Yang, Yu Qiao

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13043 2024-09-10 cs.CV cs.CL cs.LG 81%

Data Alignment for Zero-Shot Concept Generation in Dermatology AI

Soham Gadgil, Mahtab Bigverdi

专题命中 安全评测 :alignment(title);trustworthy(abstract);分类 cs.CL、cs.LG

Comments Accepted as a workshop paper to ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04641 2024-09-10 cs.LG cs.AI 81%

Stacked Universal Successor Feature Approximators for Safety in Reinforcement Learning

Ian Cannon, Washington Garcia, Thomas Gresavage, Joseph Saurine, Ian Leong, Jared Culbertson

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10311 2024-09-04 cs.CL cs.AI 81%

CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models

Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu, Meijuan An, Bikun Yang, KaiKai Zhao, Kai Wang, Shiguo Lian

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08318 2024-08-19 cs.CY cs.AI 81%

First Analysis of the EU Artifical Intelligence Act: Towards a Global Standard for Trustworthy AI?

Marion Ho-Dac

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments in French language

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04242 2024-08-09 cs.LG cs.AI cs.NE 81%

The Ungrounded Alignment Problem

Marc Pickett, Aakash Kumar Nain, Joseph Modayil, Llion Jones

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 7 pages, plus references and appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16902 2024-07-25 cs.CY cs.AI 81%

The Potential and Perils of Generative Artificial Intelligence for Quality Improvement and Patient Safety

Laleh Jalilian, Daniel McDuff, Achuta Kadambi

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12421 2024-07-18 cs.LG cs.AI 81%

SafePowerGraph: Safety-aware Evaluation of Graph Neural Networks for Transmission Power Grids

Salah Ghamizi, Aleksandar Bojchevski, Aoxiang Ma, Jun Cao

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12135 2024-07-18 cs.CY cs.AI cs.SE 81%

Trustworthy AI in practice: an analysis of practitioners' needs and challenges

Maria Teresa Baldassarre, Domenico Gigante, Marcos Kalinowski, Azzurra Ragone, Sara Tibidò

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Journal ref 28th International Conference on Evaluation and Assessment in Software Engineering (EASE 2024), June 18-21, 2024, Salerno, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16220 2024-06-25 cs.LG cs.AI cs.CV 81%

Learning Run-time Safety Monitors for Machine Learning Components

Ozan Vardal, Richard Hawkins, Colin Paterson, Chiara Picardi, Daniel Omeiza, Lars Kunze, Ibrahim Habli

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15371 2024-06-25 cs.CY cs.AI 81%

Affirmative safety: An approach to risk management for high-risk AI

Akash R. Wasil, Joshua Clymer, David Krueger, Emily Dardaman, Simeon Campos, Evan R. Murphy

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05205 2024-06-11 cs.CV cs.CL cs.LG cs.MM eess.IV 81%

CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment

Sajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Naoufel Werghi, Mohammed Bennamoun

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11195 2024-05-21 cs.LG cs.AI cs.IT math.IT 81%

Trustworthy Actionable Perturbations

Jesse Friedbaum, Sudarshan Adiga, Ravi Tandon

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments Accepted at the 41st International Conference on Machine Learning (ICML) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05466 2024-05-14 cs.CL cs.AI 81%

Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals

Joshua Clymer, Caden Juang, Severin Field

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14068 2024-04-23 cs.AI cs.LG 81%

Holistic Safety and Responsibility Evaluations of Advanced AI Models

Laura Weidinger, Joslyn Barnhart, Jenny Brennan, Christina Butterfield, Susie Young, Will Hawkins, Lisa Anne Hendricks, Ramona Comanescu, Oscar Chang, Mikel Rodriguez, Jennifer Beroshi, Dawn Bloxwich, Lev Proleev, Jilin Chen, Sebastian Farquhar, Lewis Ho, Iason Gabriel, Allan Dafoe, William Isaac

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 10 pages excluding bibliography

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10928 2024-04-16 cs.CL cs.AI 81%

FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Seonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim, Minjoon Seo

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments ICLR 2024 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06003 2024-04-10 cs.CL cs.AI 81%

FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

Zhuohao Yu, Chang Gao, Wenjin Yao, Yidong Wang, Zhengran Zeng, Wei Ye, Jindong Wang, Yue Zhang, Shikun Zhang

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments We open-source all our code at: https://github.com/WisdomShell/FreeEval

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04642 2024-04-09 cs.CL cs.AI 81%

TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction

Shuo Li, Sangdon Park, Insup Lee, Osbert Bastani

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments 23 pages, 17 figures, 2024 Annual Conference of the North American Chapter of the Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏