arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9419 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9419 篇

2107.09234 2022-03-28 cs.LG 79%

Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model Behavior

Angie Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik Strobelt

专题命中 安全评测 :alignment(title);trustworthy(abstract);分类 cs.LG

Comments 17 pages, 10 figures. Published in CHI 2022. For more details, see http://shared-interest.csail.mit.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.12653 2022-02-28 cs.LG stat.ML 79%

Bayesian autoencoders with uncertainty quantification: Towards trustworthy anomaly detection

Bang Xiang Yong, Alexandra Brintrup

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.03042 2022-02-15 cs.LG cs.CR 79%

Towards a Robust and Trustworthy Machine Learning System Development: An Engineering Perspective

Pulei Xiong, Scott Buffett, Shahrear Iqbal, Philippe Lamontagne, Mohammad Mamun, Heather Molyneaux

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments 20 pages (58 pages pre-print), 6 figures

Journal ref Journal of Information Security and Applications 65 (2022) 103121

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.05313 2022-02-14 cs.AI cs.SE 79%

Integrating Testing and Operation-related Quantitative Evidences in Assurance Cases to Argue Safety of Data-Driven AI/ML Components

Michael Kläs, Lisa Jöckel, Rasmus Adler, Jan Reich

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.01934 2022-02-07 cs.LG 79%

Smartphone-based Hard-braking Event Detection at Scale for Road Safety Services

Luyang Liu, David Racz, Kara Vaillancourt, Julie Michelman, Matt Barnes, Stefan Mellem, Paul Eastham, Bradley Green, Charles Armstrong, Rishi Bal, Shawn O'Banion, Feng Guo

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.09998 2021-12-21 cs.RO cs.LG cs.SY eess.SY 79%

Learning-based methods to model small body gravity fields for proximity operations: Safety and Robustness

Daniel Neamati, Yashwanth Kumar Nakka, Soon-Jo Chung

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments Accepted Scitech, AI for Space

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03024 2021-12-07 cs.CL 79%

Domain-oriented Language Pre-training with Adaptive Hybrid Masking and Optimal Transport Alignment

Denghui Zhang, Zixuan Yuan, Yanchi Liu, Hao Liu, Fuzhen Zhuang, Hui Xiong, Haifeng Chen

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02707 2021-10-07 cs.SE cs.AI 79%

Trustworthy Artificial Intelligence and Process Mining: Challenges and Opportunities

Andrew Pery, Majid Rafiei, Michael Simon, Wil M. P. van der Aalst

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01232 2021-10-05 cs.AI 79%

Benchmarking Safety Monitors for Image Classifiers with Machine Learning

Raul Sena Ferreira, Jean Arlat, Jeremie Guiochet, Hélène Waeselynck

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Journal ref 26th IEEE Pacific Rim International Symposium on Dependable Computing (PRDC 2021), IEEE, Dec 2021, Perth, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.10142 2021-09-28 cs.CV cs.LG 79%

Safety Metrics for Semantic Segmentation in Autonomous Driving

Chih-Hong Cheng, Alois Knoll, Hsuan-Cheng Liao

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments Paper accepted at IEEE AI Test'21

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.01287 2021-09-16 math.OC cs.LG 79%

Safety Verification and Robustness Analysis of Neural Networks via Quadratic Constraints and Semidefinite Programming

Mahyar Fazlyab, Manfred Morari, George J. Pappas

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01531 2021-09-06 cs.LG 79%

MACEst: The reliable and trustworthy Model Agnostic Confidence Estimator

Rhys Green, Matthew Rowe, Alberto Polleri

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06080 2021-08-16 cs.LG 79%

TDM: Trustworthy Decision-Making via Interpretability Enhancement

Daoming Lyu, Fangkai Yang, Hugh Kwon, Wen Dong, Levent Yilmaz, Bo Liu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Journal ref IEEE Transactions on Emerging Topics in Computational Intelligence 0 (2021) 1-12

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.08231 2021-08-13 cs.CL 79%

Word Alignment by Fine-tuning Embeddings on Parallel Corpora

Zi-Yi Dou, Graham Neubig

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments EACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.06466 2021-06-08 cs.LG stat.ML 79%

How Interpretable and Trustworthy are GAMs?

Chun-Hao Chang, Sarah Tan, Ben Lengerich, Anna Goldenberg, Rich Caruana

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments Accepted in 2021 KDD

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.00512 2021-06-02 cs.LG 79%

The Care Label Concept: A Certification Suite for Trustworthy and Resource-Aware Machine Learning

Katharina Morik, Helena Kotthaus, Lukas Heppe, Danny Heinrich, Raphael Fischer, Andreas Pauly, Nico Piatkowski

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.05434 2021-04-26 cs.SE cs.AI 79%

Developing and Operating Artificial Intelligence Models in Trustworthy Autonomous Systems

Silverio Martínez-Fernández, Xavier Franch, Andreas Jedlitschka, Marc Oriol, Adam Trendowicz

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 9 pages, 1 figure, preprint. Accepted in RCIS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.13073 2020-07-21 cs.CV cs.LG eess.IV 79%

Attributional Robustness Training using Input-Gradient Spatial Alignment

Mayank Singh, Nupur Kumari, Puneet Mangla, Abhishek Sinha, Vineeth N Balasubramanian, Balaji Krishnamurthy

专题命中 安全评测 :alignment(title);trustworthy(abstract);分类 cs.LG

Comments ECCV 2020, Code at https://github.com/nupurkmr9/Attributional-Robustness

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.01155 2020-07-01 cs.SE cs.LG 79%

Towards Probability-based Safety Verification of Systems with Components from Machine Learning

Hermann Kaindl, Stefan Kramer

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments Second (revised) version for public access

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.14750 2020-06-29 cs.CY 79%

Could regulating the creators deliver trustworthy AI?

Labhaoise Ni Fhaolain, Andrew Hines

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY

Comments To be published in The Second Workshop on Implementing Machine Ethics, Dublin, Ireland, 30 June 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.07213 2020-04-22 cs.CY 79%

Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims

Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, Tegan Maharaj, Pang Wei Koh, Sara Hooker, Jade Leung, Andrew Trask, Emma Bluemke, Jonathan Lebensold, Cullen O'Keefe, Mark Koren, Théo Ryffel, JB Rubinovitz, Tamay Besiroglu, Federica Carugati, Jack Clark, Peter Eckersley, Sarah de Haas, Maritza Johnson, Ben Laurie, Alex Ingerman, Igor Krawczuk, Amanda Askell, Rosario Cammarota, Andrew Lohn, David Krueger, Charlotte Stix, Peter Henderson, Logan Graham, Carina Prunkl, Bianca Martin, Elizabeth Seger, Noa Zilberman, Seán Ó hÉigeartaigh, Frens Kroeger, Girish Sastry, Rebecca Kagan, Adrian Weller, Brian Tse, Elizabeth Barnes, Allan Dafoe, Paul Scharre, Ariel Herbert-Voss, Martijn Rasser, Shagun Sodhani, Carrick Flynn, Thomas Krendl Gilbert, Lisa Dyer, Saif Khan, Yoshua Bengio, Markus Anderljung

专题命中 安全评测 :trustworthy(title);safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.08210 2020-02-20 cs.AI cs.CR cs.RO 79%

A Structured Approach to Trustworthy Autonomous/Cognitive Systems

Henrik J. Putzer, Ernest Wozniak

专题命中 安全评测 :trustworthy(title);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.06276 2020-02-18 cs.AI 79%

Trustworthy AI

Jeannette M. Wing

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.03515 2019-10-09 cs.AI cs.HC 79%

Designing Trustworthy AI: A Human-Machine Teaming Framework to Guide Development

Carol J. Smith

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments Presented at AAAI FSS-19: Artificial Intelligence in Government and Public Sector, Arlington, Virginia, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.10768 2019-09-19 cs.AI 79%

Deep Trustworthy Knowledge Tracing

Heonseok Ha, Uiwon Hwang, Yongjun Hong, Jahee Jang, Sungroh Yoon

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.06434 2016-10-21 stat.ML cs.LG 79%

Kernel Alignment for Unsupervised Transfer Learning

Ievgen Redko, Younès Bennani

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1109.6029 2011-09-29 cs.AI 79%

An Improved Search Algorithm for Optimal Multiple-Sequence Alignment

S. Schroedl

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Journal ref Journal Of Artificial Intelligence Research, Volume 23, pages 587-623, 2005

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25809 2026-06-25 cs.HC 新提交 79%

Designing Trustworthy LLM-based Wellbeing Recommendation through Controllable Interaction

通过可控交互设计可信赖的基于LLM的健康推荐

Alan Said, Alexandra Weilenmann

专题命中 安全评测 :trustworthy(title,comments);alignment(abstract)

AI总结 提出通过显式交互约束(如指导策略、解释风格、直接性程度和用户控制机制)构建系统级框架,使LLM推荐在保持适应性的同时透明、可控且与人类健康对齐。

Comments Accepted to the 1st Workshop on Trustworthy and Adaptive LLMs for Mental and Physical Wellbeing in Recommendations @UMAP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18324 2024-05-29 cs.RO 79%

Value Alignment and Trust in Human-Robot Interaction: Insights from Simulation and User Study

Shreyas Bhat, Joseph B. Lyons, Cong Shi, X. Jessie Yang

专题命中 安全评测 :alignment(title,abstract)

Comments This is a preprint of the following chapter: Bhat et al., Value Alignment and Trust in Human-Robot Interaction: Insights from Simulation and User Study, published in "Emerging Frontiers in Human-Robot Interaction", edited by Ramana Kumar Vinjamuri, 2024, Springer Nature reproduced with permission of Springer Nature. The final authenticated version is available online at: [INSERT LINK HERE]

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.07619 2021-06-28 cs.MA 79%

Value Alignment Equilibrium in Multiagent Systems

Nieves Montes, Carles Sierra

专题命中 安全评测 :alignment(title,abstract);trustworthy(journal_ref)

Comments 1st TAILOR Workshop at ECAI 2020

Journal ref In: Trustworthy AI - Integrating Learning, Optimization and Reasoning. TAILOR 2020. Lecture Notes in Computer Science, vol 12641. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏