arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2304.00133 2024-04-19 cs.LG cs.HC 57%

DeforestVis: Behavior Analysis of Machine Learning Models with Surrogate Decision Stumps

Angelos Chatzimparmpas, Rafael M. Martins, Alexandru C. Telea, Andreas Kerren

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments This manuscript is accepted for publication in Computer Graphics Forum (CGF)

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.11737 2024-04-19 cs.LG cs.HC stat.ML 57%

The State of the Art in Enhancing Trust in Machine Learning Models with the Use of Visualizations

A. Chatzimparmpas, R. Martins, I. Jusufi, K. Kucher, Fabrice Rossi, A. Kerren

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Journal ref Computer Graphics Forum 2020, 39(3), 713-756

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00334 2024-04-19 cs.LG cs.HC stat.ML 57%

VisRuler: Visual Analytics for Extracting Decision Rules from Bagged and Boosted Decision Trees

Angelos Chatzimparmpas, Rafael M. Martins, Andreas Kerren

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments This manuscript is accepted for publication in the Information Visualization (IV) - SAGE Journals

Journal ref Information Visualization, 2021, 22(2), 115-139

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10782 2024-04-18 cs.CR cs.AI 57%

Quantifying AI Vulnerabilities: A Synthesis of Complexity, Dynamical Systems, and Game Theory

B Kereopa-Yorke

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11060 2024-04-16 quant-ph cs.LG stat.ML 57%

Exponential concentration in quantum kernel methods

Supanut Thanasilp, Samson Wang, M. Cerezo, Zoë Holmes

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 15+50 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08069 2024-04-15 cs.LG 57%

Persistent Classification: A New Approach to Stability of Data and Adversarial Examples

Brian Bell, Michael Geyer, David Glickenstein, Keaton Hamm, Carlos Scheidegger, Amanda Fernandez, Juston Moore

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07063 2024-04-11 cs.RO cs.AI 57%

LaPlaSS: Latent Space Planning for Stochastic Systems

Marlyse Reeves, Brian C. Williams

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04442 2024-04-10 cs.AI 57%

Exploring Autonomous Agents through the Lens of Large Language Models: A Review

Saikat Barua

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 47 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03769 2024-04-08 cs.SE cs.LG cs.SY eess.SY 57%

On Extending the Automatic Test Markup Language (ATML) for Machine Learning

Tyler Cody, Bingtong Li, Peter A. Beling

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Accepted by the 18th Annual IEEE International Systems Conference (SysCon)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03036 2024-04-05 cs.CL 57%

MuLan: A Study of Fact Mutability in Language Models

Constanza Fierro, Nicolas Garneau, Emanuele Bugliarello, Yova Kementchedjhieva, Anders Søgaard

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13332 2024-04-05 cs.LG stat.ME 57%

Causal hybrid modeling with double machine learning

Kai-Hendrik Cohrs, Gherardo Varando, Nuno Carvalhais, Markus Reichstein, Gustau Camps-Valls

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09390 2024-04-05 cs.CL 57%

LEEETs-Dial: Linguistic Entrainment in End-to-End Task-oriented Dialogue systems

Nalin Kumar, Ondřej Dušek

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to NAACL Findings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01667 2024-04-03 cs.CL 57%

METAL: Towards Multilingual Meta-Evaluation

Rishav Hada, Varun Gumma, Mohamed Ahmed, Kalika Bali, Sunayana Sitaram

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to NAACL 2024 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09648 2024-04-03 cs.CL 57%

Event Causality Is Key to Computational Story Understanding

Yidan Sun, Qin Chao, Boyang Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00205 2024-04-02 cs.CL 57%

Conceptual and Unbiased Reasoning in Language Models

Ben Zhou, Hongming Zhang, Sihao Chen, Dian Yu, Hongwei Wang, Baolin Peng, Dan Roth, Dong Yu

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Preprint under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01327 2024-04-02 cs.CV cs.AI cs.MM 57%

Can I Trust Your Answer? Visually Grounded Video Question Answering

Junbin Xiao, Angela Yao, Yicong Li, Tat Seng Chua

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted to CVPR'24. (Compared with preprint version, we mainly improve the presentation, discuss more related works, and extend experiments in Appendix.)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.20078 2024-04-01 cs.CV cs.LG 57%

Negative Label Guided OOD Detection with Pretrained Vision-Language Models

Xue Jiang, Feng Liu, Zhen Fang, Hong Chen, Tongliang Liu, Feng Zheng, Bo Han

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments ICLR 2024 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11128 2024-03-28 cs.CL 57%

Beyond Static Evaluation: A Dynamic Approach to Assessing AI Assistants' API Invocation Capabilities

Honglin Mu, Yang Xu, Yunlong Feng, Xiaofeng Han, Yitong Li, Yutai Hou, Wanxiang Che

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted at LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16858 2024-03-26 cs.AI 57%

XAIport: A Service Framework for the Early Adoption of XAI in AI Model Development

Zerui Wang, Yan Liu, Abishek Arumugam Thiruselvi, Abdelwahab Hamou-Lhadj

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted at the ICSE'24 conference, NIER track

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16609 2024-03-26 cs.CL 57%

Conversational Grounding: Annotation and Analysis of Grounding Acts and Grounding Units

Biswesh Mohapatra, Seemab Hassan, Laurent Romary, Justine Cassell

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Journal ref LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09962 2024-03-26 cs.LG stat.ML 57%

Distributional Robustness Bounds Generalization Errors

Shixiong Wang, Haowei Wang

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Updated Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14564 2024-03-22 cs.CL 57%

Language Models Hallucinate, but May Excel at Fact Verification

Jian Guan, Jesse Dodge, David Wadden, Minlie Huang, Hao Peng

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Accepted in NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12371 2024-03-20 cs.LG 57%

Advancing Time Series Classification with Multimodal Language Modeling

Mingyue Cheng, Yiheng Chen, Qi Liu, Zhiding Liu, Yucong Luo

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11844 2024-03-19 cs.LG eess.SP math.OC 57%

Near-Optimal Solutions of Constrained Learning Problems

Juan Elenter, Luiz F. O. Chamon, Alejandro Ribeiro

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14396 2024-03-18 cs.SE cs.LG cs.PL 57%

Guess & Sketch: Language Model Guided Transpilation

Celine Lee, Abdulrahman Mahmoud, Michal Kurek, Simone Campanoni, David Brooks, Stephen Chong, Gu-Yeon Wei, Alexander M. Rush

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09491 2024-03-15 cs.LG math.DS 57%

On using Machine Learning Algorithms for Motorcycle Collision Detection

Philipp Rodegast, Steffen Maier, Jonas Kneifl, Jörg Fehr

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08880 2024-03-15 cs.LG 57%

REFRESH: Responsible and Efficient Feature Reselection Guided by SHAP Values

Shubham Sharma, Sanghamitra Dutta, Emanuele Albini, Freddy Lecue, Daniele Magazzeni, Manuela Veloso

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02535 2024-03-14 cs.CV cs.LG 57%

Learning to Generate Training Datasets for Robust Semantic Segmentation

Marwane Hariat, Olivier Laurent, Rémi Kazmierczak, Shihao Zhang, Andrei Bursuc, Angela Yao, Gianni Franchi

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Published as a conference paper at WACV 2024

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2024. p. 3894-3905

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07321 2024-03-13 cs.CL 57%

GPT-generated Text Detection: Benchmark Dataset and Tensor-based Detection Method

Zubair Qazi, William Shiao, Evangelos E. Papalexakis

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 4 pages, 2 figures, published in the WWW 2024 Short Papers Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05583 2024-03-12 cs.HC cs.AI cs.SD eess.AS 57%

A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition

Tyler Benster, Guy Wilson, Reshef Elisha, Francis R Willett, Shaul Druckmann

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏