arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2401.07529 2024-06-04 cs.CV cs.CL 57%

MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception

Yuhao Wang, Yusheng Liao, Heyang Liu, Hongcheng Liu, Yu Wang, Yanfeng Wang

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20991 2024-06-03 cs.CV cs.LG 57%

Hard Cases Detection in Motion Prediction by Vision-Language Foundation Models

Yi Yang, Qingwen Zhang, Kei Ikemura, Nazre Batool, John Folkesson

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments IEEE Intelligent Vehicles Symposium (IV) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12424 2024-06-03 cs.RO cs.LG 57%

Rethinking Robustness Assessment: Adversarial Attacks on Learning-based Quadrupedal Locomotion Controllers

Fan Shi, Chong Zhang, Takahiro Miki, Joonho Lee, Marco Hutter, Stelian Coros

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments RSS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19041 2024-05-30 cs.CL cs.SD eess.AS 57%

BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation

Chen Wang, Minpeng Liao, Zhongqiang Huang, Jiajun Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11762 2024-05-30 cs.LG 57%

Interpretability of Statistical, Machine Learning, and Deep Learning Models for Landslide Susceptibility Mapping in Three Gorges Reservoir Area

Cheng Chen, Lei Fan

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15777 2024-05-30 cs.CL 57%

A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry

Yining Huang, Keke Tang, Meilian Chen, Boyuan Wang

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 42 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15307 2024-05-27 cs.CL 57%

Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation

Ge Qu, Jinyang Li, Bowen Li, Bowen Qin, Nan Huo, Chenhao Ma, Reynold Cheng

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to ACL Findings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05634 2024-05-24 cs.CL 57%

Towards Verifiable Generation: A Benchmark for Knowledge-aware Language Model Attribution

Xinze Li, Yixin Cao, Liangming Pan, Yubo Ma, Aixin Sun

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments acl findings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13191 2024-05-24 cs.LG 57%

Pragmatic auditing: a pilot-driven approach for auditing Machine Learning systems

Djalel Benbouzid, Christiane Plociennik, Laura Lucaj, Mihai Maftei, Iris Merget, Aljoscha Burchardt, Marc P. Hauer, Abdeldjallil Naceri, Patrick van der Smagt

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14459 2024-05-22 cs.SE cs.AI 57%

LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations

Rebeka Tóth, Tamas Bisztray, László Erdodi

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16457 2024-05-21 cs.LG eess.SP q-bio.NC 57%

SI-SD: Sleep Interpreter through awake-guided cross-subject Semantic Decoding

Hui Zheng, Zhong-Tao Chen, Hai-Teng Wang, Jian-Yang Zhou, Lin Zheng, Pei-Yang Lin, Yun-Zhe Liu

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10647 2024-05-17 cs.CL 57%

Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models

Rima Hazra, Sayan Layek, Somnath Banerjee, Soujanya Poria

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Accepted at ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08035 2024-05-15 cs.HC cs.AI 57%

A LLM-based Controllable, Scalable, Human-Involved User Simulator Framework for Conversational Recommender Systems

Lixi Zhu, Xiaowen Huang, Jitao Sang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07603 2024-05-14 cs.RO cs.AI 57%

Reducing Risk for Assistive Reinforcement Learning Policies with Diffusion Models

Andrii Tytarenko

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.04122 2024-05-13 cs.HC cs.AI 57%

From Prompt Engineering to Prompt Science With Human in the Loop

Chirag Shah

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11937 2024-05-10 cs.AI 57%

On the Definition of Appropriate Trust and the Tools that Come with it

Helena Löfström

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 8 pages, 3 figures, Conference: ICDATA 2023

Journal ref 2023 Congress in Computer Science, Computer Engineering, & Applied Computing (CSCE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12545 2024-05-08 cs.CL 57%

TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness

Danna Zheng, Danyang Liu, Mirella Lapata, Jeff Z. Pan

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02583 2024-05-07 cs.AI 57%

Explainable Interface for Human-Autonomy Teaming: A Survey

Xiangqi Kong, Yang Xing, Antonios Tsourdos, Ziyue Wang, Weisi Guo, Adolfo Perrusquia, Andreas Wikander

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 45 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14658 2024-05-07 cs.CL 57%

Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response

Yongkang Liu, Shi Feng, Daling Wang, Yifei Zhang, Hinrich Schütze

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01156 2024-05-03 cs.CV cs.AI 57%

Self-Supervised Learning for Interventional Image Analytics: Towards Robust Device Trackers

Saahil Islam, Venkatesh N. Murthy, Dominik Neumann, Badhan Kumar Das, Puneet Sharma, Andreas Maier, Dorin Comaniciu, Florin C. Ghesu

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18663 2024-04-30 cs.CV cs.LG cs.RO cs.SE 57%

Terrain characterisation for online adaptability of automated sonar processing: Lessons learnt from operationally applying ATR to sidescan sonar in MCM applications

Thomas Guerneve, Stephanos Loizou, Andrea Munafo, Pierre-Yves Mignotte

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Presented at UACE (Underwater Acoustics Conference & Exhibition) 2023, Kalamata, Greece

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16681 2024-04-30 cs.CV cs.AI 57%

Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations

Maximilian Dreyer, Reduan Achtibat, Wojciech Samek, Sebastian Lapuschkin

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 39 pages (8 pages manuscript, 3 pages references, 28 pages appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18183 2024-04-30 q-fin.RM cs.AI 57%

Innovative Application of Artificial Intelligence Technology in Bank Credit Risk Management

Shuochen Bi, Wenqing Bao

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 6 pages, 1 figure, 2 tables

Journal ref International Journal of Global Economics and Management ISSN: 3005-9690 (Print), ISSN: 3005-8090 (Online) | Volume 2, Number 3, Year 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17350 2024-04-29 cs.LG cs.CV cs.MA 57%

On the Road to Clarity: Exploring Explainable AI for World Models in a Driver Assistance System

Mohamed Roshdi, Julian Petzold, Mostafa Wahby, Hussein Ebrahim, Mladen Berekovic, Heiko Hamann

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 8 pages, 6 figures, to be published in IEEE CAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08715 2024-04-29 cs.CL 57%

SOTOPIA-$π$: Interactive Learning of Socially Intelligent Language Agents

Ruiyi Wang, Haofei Yu, Wenxin Zhang, Zhengyang Qi, Maarten Sap, Graham Neubig, Yonatan Bisk, Hao Zhu

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03314 2024-04-29 eess.SY cs.LG cs.MA cs.RO cs.SY 57%

Collision Avoidance Verification of Multiagent Systems with Learned Policies

Zihao Dong, Shayegan Omidshafiei, Michael Everett

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14544 2024-04-24 cs.CL 57%

WangLab at MEDIQA-CORR 2024: Optimized LLM-based Programs for Medical Error Detection and Correction

Augustin Toma, Ronald Xie, Steven Palayew, Patrick R. Lawler, Bo Wang

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12653 2024-04-22 cs.AI 57%

How Real Is Real? A Human Evaluation Framework for Unrestricted Adversarial Examples

Dren Fazlija, Arkadij Orlov, Johanna Schrader, Monty-Maximilian Zühlke, Michael Rohs, Daniel Kudenko

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 3 pages, 3 figures, AAAI 2024 Spring Symposium on User-Aligned Assessment of Adaptive AI Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12272 2024-04-19 cs.HC cs.AI 57%

Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences

Shreya Shankar, J. D. Zamfirescu-Pereira, Björn Hartmann, Aditya G. Parameswaran, Ian Arawjo

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 16 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12005 2024-04-19 cs.HC cs.LG stat.ML 57%

Visualization for Trust in Machine Learning Revisited: The State of the Field in 2023

Angelos Chatzimparmpas, Kostiantyn Kucher, Andreas Kerren

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments This manuscript is accepted for publication in the IEEE Computer Graphics and Applications Journal (IEEE CG&A)

详情

展开后加载摘要…

URL PDF HTML 收藏