arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9419 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9419 篇

2403.06128 2024-03-12 eess.IV cs.CV 78%

Low-dose CT Denoising with Language-engaged Dual-space Alignment

Zhihao Chen, Tao Chen, Chenhui Wang, Chuang Niu, Ge Wang, Hongming Shan

专题命中 安全评测 :alignment(title,abstract)

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05135 2024-03-11 cs.CV 78%

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, Gang Yu

专题命中 安全评测 :alignment(title,abstract)

Comments Project Page: https://ella-diffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12515 2024-02-23 cs.IR 78%

A Survey on Trustworthy Recommender Systems

Yingqiang Ge, Shuchang Liu, Zuohui Fu, Juntao Tan, Zelong Li, Shuyuan Xu, Yunqi Li, Yikun Xian, Yongfeng Zhang

专题命中 安全评测 :trustworthy(title,abstract)

Comments Accepted by ACM Transactions on Recommender Systems (TORS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13232 2024-02-21 cs.CV cs.RO 78%

A Touch, Vision, and Language Dataset for Multimodal Alignment

Letian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch, Jaimyn Drake, Joseph Ortiz, Mustafa Mukadam, Mike Lambeta, Roberto Calandra, Ken Goldberg

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05893 2024-02-09 cs.HC 78%

Personalizing Driver Safety Interfaces via Driver Cognitive Factors Inference

Emily S Sumner, Jonathan DeCastro, Jean Costa, Deepak E Gopinath, Everlyne Kimani, Shabnam Hakimi, Allison Morgan, Andrew Best, Hieu Nguyen, Daniel J Brooks, Bassam ul Haq, Andrew Patrikalakis, Hiroshi Yasuda, Kate Sieck, Avinash Balachandran, Tiffany Chen, Guy Rosman

专题命中 安全评测 :safety(title,abstract)

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.14185 2024-01-23 cs.CR 78%

MALIGN: Explainable Static Raw-byte Based Malware Family Classification using Sequence Alignment

Shoumik Saha, Sadia Afroz, Atif Rahman

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00991 2024-01-03 cs.CR 78%

A Novel Evaluation Framework for Assessing Resilience Against Prompt Injection Attacks in Large Language Models

Daniel Wankit Yip, Aysan Esmradi, Chun Fai Chan

专题命中 安全评测 :prompt injection(title,abstract)

Comments Accepted to be published in the Proceedings of The 10th IEEE CSDE 2023, the Asia-Pacific Conference on Computer Science and Data Engineering 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02310 2023-12-06 cs.CV cs.AI cs.CL cs.LG 78%

VaQuitA: Enhancing Alignment in LLM-Assisted Video Understanding

Yizhou Wang, Ruiyi Zhang, Haoliang Wang, Uttaran Bhattacharya, Yun Fu, Gang Wu

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17945 2023-12-01 cs.CV 78%

Contrastive Vision-Language Alignment Makes Efficient Instruction Learner

Lizhao Liu, Xinyu Sun, Tianhang Xiang, Zhuangwei Zhuang, Liuren Yin, Mingkui Tan

专题命中 安全评测 :alignment(title,abstract)

Comments 17 pages, 10 pages for main paper, 7 pages for supplementary

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.15367 2023-11-29 stat.AP stat.ME 78%

Targeted learning in observational studies with multi-valued treatments: An evaluation of antipsychotic drug treatment safety

Jason Poulos, Marcela Horvitz-Lennon, Katya Zelevinsky, Tudor Cristea-Platon, Thomas Huijskens, Pooja Tyagi, Jiaju Yan, Jordi Diaz, Sharon-Lise Normand

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07438 2023-10-12 cs.CV cs.RO 78%

DESTINE: Dynamic Goal Queries with Temporal Transductive Alignment for Trajectory Prediction

Rezaul Karim, Soheil Mohamad Alizadeh Shabestary, Amir Rasouli

专题命中 安全评测 :alignment(title,abstract)

Comments 6 tables 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14115 2023-09-12 cs.LG cs.AI cs.CL stat.ME stat.ML 78%

Towards Trustworthy Explanation: On Causal Rationalization

Wenbo Zhang, Tong Wu, Yunlong Wang, Yong Cai, Hengrui Cai

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI、cs.LG

Comments In Proceedings of the 40th International Conference on Machine Learning (ICML) GitHub Repository: https://github.com/onepounchman/Causal-Retionalization

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00920 2023-09-06 cs.MA 78%

Trustworthy Distributed Average Consensus based on Locally Assessed Trust Evaluations

Christoforos N. Hadjicostis, Alejandro D. Dominguez-Garcia

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00249 2023-09-04 cs.RO 78%

Suicidal Pedestrian: Generation of Safety-Critical Scenarios for Autonomous Vehicles

Yuhang Yang, Kalle Kujanpaa, Amin Babadi, Joni Pajarinen, Alexander Ilin

专题命中 安全评测 :safety(title,abstract)

Comments 6 pages; 5 figures; 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.02215 2023-06-22 cs.RO 78%

A Survey on Safety-Critical Driving Scenario Generation -- A Methodological Perspective

Wenhao Ding, Chejian Xu, Mansur Arief, Haohong Lin, Bo Li, Ding Zhao

专题命中 安全评测 :safety(title,abstract)

Comments 18 pages, 5 figures. IEEE Transactions on Intelligent Transportation Systems (T-ITS) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11291 2023-05-22 cs.HC 78%

A systematic review of safety-critical scenarios between automated vehicles and vulnerable road users

Aditya Deshmukh, Zifei Wang, Aaron Gunn, Huizhong Guo, Rini Sherony, Fred Feng, Brian Lin, Shan Bao, Feng Zhou

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04514 2023-04-11 cs.CV 78%

DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment

Lewei Yao, Jianhua Han, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, Hang Xu

专题命中 安全评测 :alignment(title,abstract)

Comments Accepted to CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.04986 2023-03-10 cs.HC 78%

Work with AI and Work for AI: Autonomous Vehicle Safety Drivers' Lived Experiences

Mengdi Chu, Keyu Zong, Xin Shu, Jiangtao Gong, Zicong Lu, Kaimin Guo, Xinyi Dai, Guyue Zhou

专题命中 安全评测 :safety(title,abstract)

Comments 17 pages, 2 figures

Journal ref CHI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.02635 2023-03-07 cs.CV cs.MM 78%

VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media Reasoning

Kang Chen, Xiangqian Wu

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.06613 2023-02-15 eess.IV 78%

Explainable artificial intelligence toward usable and trustworthy computer-aided early diagnosis of multiple sclerosis from Optical Coherence Tomography

Monica Hernandez, Ubaldo Ramon-Julvez, Elisa Vilades, Beatriz Cordon, Elvira Mayordomo, Elena Garcia-Martin

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11136 2023-01-27 stat.ML 78%

Conformal Prediction for Trustworthy Detection of Railway Signals

Léo Andéol, Thomas Fel, Florence De Grancey, Luca Mossina

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.01599 2022-11-11 cs.HC 78%

A Gaze Data-based Comparative Study to Build a Trustworthy Human-AI Collaboration in Crash Anticipation

Yu Li, Muhammad Monjurul Karim, Ruwen Qin

专题命中 安全评测 :trustworthy(title);safety(abstract)

Comments Revised and submitted to International Conference on Transportation and Development (ICTD 2023) on Nov 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.17291 2022-11-01 cs.IT math.IT 78%

SIX-Trust for 6G: Towards a Secure and Trustworthy 6G Network

Yiying Wang, Xin Kang, Tieyan Li, Haiguang Wang, Cheng-Kang Chu, Zhongding Lei

专题命中 安全评测 :trustworthy(title,abstract)

Comments 7 pages, 3 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09682 2022-11-01 cs.RO 78%

SafeBench: A Benchmarking Platform for Safety Evaluation of Autonomous Vehicles

Chejian Xu, Wenhao Ding, Weijie Lyu, Zuxin Liu, Shuai Wang, Yihan He, Hanjiang Hu, Ding Zhao, Bo Li

专题命中 安全评测 :safety(title,abstract)

Comments Published as a conference paper at NeurIPS 2022 (Track on Datasets and Benchmarks)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.07198 2022-10-14 cs.CL cs.AI cs.LG 78%

Towards Trustworthy Automatic Diagnosis Systems by Emulating Doctors' Reasoning with Deep Reinforcement Learning

Arsene Fansi Tchango, Rishab Goel, Julien Martel, Zhi Wen, Gaetan Marceau Caron, Joumana Ghosn

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI、cs.LG

Comments Camera ready. NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.02403 2022-09-07 cs.HC 78%

Guidelines to Develop Trustworthy Conversational Agents for Children

Marina Escobar-Planas, Emilia Gómez, Carlos-D Martínez-Hinarejos

专题命中 安全评测 :trustworthy(title,abstract)

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03083 2022-07-26 cs.CV cs.RO 78%

Construction Site Safety Monitoring and Excavator Activity Analysis System

Sibo Zhang, Liangjun Zhang

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.15757 2022-06-01 cs.DC cs.CR 78%

Dropbear: Machine Learning Marketplaces made Trustworthy with Byzantine Model Agreement

Alex Shamis, Peter Pietzuch, Antoine Delignat-Lavaud, Andrew Paverd, Manuel Costa

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00688 2022-04-05 cs.HC 78%

Designing AI for Online-to-Offline Safety Risks with Young Women: The Context of Social Matching

Douglas Zytko, Hanan Aljasim

专题命中 安全评测 :safety(title,abstract)

Comments Accepted to the ACM CSCW 2021 workshop "MOSafely: Building an Open-Source HCAI Community to Make the Internet a Safer Place For Youth"

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04983 2021-10-14 eess.SY cs.SY 78%

Understanding the Safety Requirements for Learning-based Power Systems Operations

Yize Chen, Daniel Arnold, Yuanyuan Shi, Sean Peisert

专题命中 安全评测 :safety(title,abstract)

Comments In submission

详情

展开后加载摘要…

URL PDF HTML 收藏