arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9400 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9400 篇

2412.18971 2024-12-30 cs.LG 74%

Adopting Trustworthy AI for Sleep Disorder Prediction: Deep Time Series Analysis with Temporal Attention Mechanism and Counterfactual Explanations

Pegah Ahadian, Wei Xu, Sherry Wang, Qiang Guan

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Journal ref IEEE Bigdata 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17527 2024-12-24 cs.AI 74%

Enhancing Cancer Diagnosis with Explainable & Trustworthy Deep Learning Models

Badaru I. Olumuyiwa, The Anh Han, Zia U. Shamszaman

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11713 2024-12-17 cs.CL cs.SE 74%

Seeker: Towards Exception Safety Code Generation with Intermediate Language Agents Framework

Xuanming Zhang, Yuxuan Chen, Yiming Zheng, Zhexin Zhang, Yuan Yuan, Minlie Huang

专题命中 安全评测 :safety(title);分类 cs.CL

Comments 30 pages, 9 figures, submitted to ARR Dec

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18222 2024-12-09 cs.AI 74%

Trustworthy AI: Securing Sensitive Data in Large Language Models

Georgios Feretzakis, Vassilios S. Verykios

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments 40 pages, 1 figure

Journal ref AI 5(4), 2773-2800 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08889 2024-11-15 cs.HC cs.AI cs.SD eess.AS 74%

Multilingual Standalone Trustworthy Voice-Based Social Network for Disaster Situations

Majid Behravan, Elham Mohammadrezaei, Mohamed Azab, Denis Gracanin

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments Accepted for publication in IEEE UEMCON 2024, to appear in December 2024. 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03376 2024-11-07 cs.DC cs.AI cs.NI 74%

An Open API Architecture to Discover the Trustworthy Explanation of Cloud AI Services

Zerui Wang, Yan Liu, Jun Huang

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments Published in: IEEE Transactions on Cloud Computing ( Volume: 12, Issue: 2, April-June 2024)

Journal ref IEEE Transactions on Cloud Computing ( Volume: 12, Issue: 2, April-June 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01335 2024-11-04 cs.CV cs.AI 74%

BehAVE: Behaviour Alignment of Video Game Encodings

Nemanja Rašajski, Chintan Trivedi, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis

专题命中 安全评测 :alignment(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15706 2024-09-25 cs.HC cs.AI 74%

Improving Emotional Support Delivery in Text-Based Community Safety Reporting Using Large Language Models

Yiren Liu, Yerong Li, Ryan Mayfield, Yun Huang

专题命中 安全评测 :safety(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04407 2024-07-08 cs.LG 74%

Trustworthy Classification through Rank-Based Conformal Prediction Sets

Rui Luo, Zhixin Zhou

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01862 2024-07-02 cs.CL cs.DB 74%

$R^3$-NL2GQL: A Model Coordination and Knowledge Graph Alignment Approach for NL2GQL

Yuhang Zhou, Yu He, Siyu Tian, Yuchen Ni, Zhangyue Yin, Xiang Liu, Chuanjun Ji, Sen Liu, Xipeng Qiu, Guangnan Ye, Hongfeng Chai

专题命中 安全评测 :alignment(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11845 2024-06-21 cs.CY cs.HC 74%

Decoding the Digital Fine Print: Navigating the potholes in Terms of service/ use of GenAI tools against the emerging need for Transparent and Trustworthy Tech Futures

Sundaraparipurnan Narayanan

专题命中 安全评测 :trustworthy(title);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07820 2024-06-13 cs.CV cs.LG 74%

Are Objective Explanatory Evaluation metrics Trustworthy? An Adversarial Analysis

Prithwijit Chowdhury, Mohit Prabhushankar, Ghassan AlRegib, Mohamed Deriche

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07080 2024-06-12 cs.CL 74%

DARA: Decomposition-Alignment-Reasoning Autonomous Language Agent for Question Answering over Knowledge Graphs

Haishuo Fang, Xiaodan Zhu, Iryna Gurevych

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments Accepted by ACL2024 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11698 2024-04-19 cs.AI cs.DC 74%

A Secure and Trustworthy Network Architecture for Federated Learning Healthcare Applications

Antonio Boiano, Marco Di Gennaro, Luca Barbieri, Michele Carminati, Monica Nicoli, Alessandro Redondi, Stefano Savazzi, Albert Sund Aillet, Diogo Reis Santos, Luigi Serio

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09668 2024-03-18 cs.CV cs.AI 74%

Trustworthy Automated Driving through Qualitative Scene Understanding and Explanations

Nassim Belmecheri, Arnaud Gotlieb, Nadjib Lazaar, Helge Spieker

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments Transport Research Arena (TRA) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00446 2024-02-02 cs.CL 74%

Improving Dialog Safety using Socially Aware Contrastive Learning

Souvik Das, Rohini K. Srihari

专题命中 安全评测 :safety(title);分类 cs.CL

Comments SCI-CHAT@EACL2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04231 2023-12-08 cs.CV cs.AI 74%

Adventures of Trustworthy Vision-Language Models: A Survey

Mayank Vatsa, Anubhooti Jain, Richa Singh

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments Accepted in AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01213 2023-12-05 cs.AR cs.AI 74%

Recent Advances in Scalable Energy-Efficient and Trustworthy Spiking Neural networks: from Algorithms to Technology

Souvik Kundu, Rui-Jie Zhu, Akhilesh Jaiswal, Peter A. Beerel

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16822 2023-10-24 cs.LG cs.DC cs.SE 74%

Rethinking Certification for Trustworthy Machine Learning-Based Applications

Marco Anisetti, Claudio A. Ardagna, Nicola Bena, Ernesto Damiani

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Accepted in IEEE Internet Computing; 6 pages, 1 figure, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10147 2023-07-26 cs.IT cs.LG math.IT 74%

TEFL: Turbo Explainable Federated Learning for 6G Trustworthy Zero-Touch Network Slicing

Swastika Roy, Hatim Chergui, Christos Verikoukis

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Overlapes with the new version

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01902 2023-05-23 cs.LG stat.ML 74%

Barycentric-alignment and reconstruction loss minimization for domain generalization

Boyang Lyu, Thuan Nguyen, Prakash Ishwar, Matthias Scheutz, Shuchin Aeron

专题命中 安全评测 :alignment(title);分类 cs.LG

Comments This article has been accepted for publication in IEEE Access

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11597 2023-05-22 cs.AI 74%

Flexible and Inherently Comprehensible Knowledge Representation for Data-Efficient Learning and Trustworthy Human-Machine Teaming in Manufacturing Environments

Vedran Galetić, Alistair Nottle

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.01204 2023-04-05 cs.AI 74%

Automatic Geo-alignment of Artwork in Children's Story Books

Jakub J. Dylag, Victor Suarez, James Wald, Aneesha Amodini Uvara

专题命中 安全评测 :alignment(title);分类 cs.AI

Comments Master's project

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06045 2022-12-13 cs.LG 74%

PERFEX: Classifier Performance Explanations for Trustworthy AI Systems

Erwin Walraven, Ajaya Adhikari, Cor J. Veenman

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00572 2022-10-04 cs.CL 74%

Risk-graded Safety for Handling Medical Queries in Conversational AI

Gavin Abercrombie, Verena Rieser

专题命中 安全评测 :safety(title);分类 cs.CL

Comments Accepted for publication at AACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05257 2022-09-13 cs.HC cs.LG 74%

TruVR: Trustworthy Cybersickness Detection using Explainable Machine Learning

Ripan Kumar Kundu, Rifatul Islam, Prasad Calyam, Khaza Anuarul Hoque

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Accepted copy, to be published in ISAMR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.04608 2022-08-10 cs.IR cs.AI 74%

Using Sentence Embeddings and Semantic Similarity for Seeking Consensus when Assessing Trustworthy AI

Dennis Vetter, Jesmin Jahan Tithi, Magnus Westerlund, Roberto V. Zicari, Gemma Roig

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12630 2022-05-26 cs.CL cs.CV 74%

Multimodal Knowledge Alignment with Reinforcement Learning

Youngjae Yu, Jiwan Chung, Heeseung Yun, Jack Hessel, JaeSung Park, Ximing Lu, Prithviraj Ammanabrolu, Rowan Zellers, Ronan Le Bras, Gunhee Kim, Yejin Choi

专题命中 安全评测 :alignment(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.04816 2022-05-13 cs.CR cs.LG 74%

Towards a trustworthy, secure and reliable enclave for machine learning in a hospital setting: The Essen Medical Computing Platform (EMCP)

Hendrik F. R. Schmidt, Jörg Schlötterer, Marcel Bargull, Enrico Nasca, Ryan Aydelott, Christin Seifert, Folker Meyer

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments 9 pages, 5 figures, to be published in the proceedings of the 2021 IEEE CogMI conference. Christin Seifert and Folker Meyer are co-senior authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.15112 2021-10-01 cs.LG cs.CR 74%

Interpretability in Safety-Critical FinancialTrading Systems

Gabriel Deza, Adelin Travers, Colin Rowat, Nicolas Papernot

专题命中 安全评测 :safety(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏