arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-29 至 2025-07-29 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 4 篇

2506.12088 2025-07-29 cs.CR cs.CY 86%

Risks & Benefits of LLMs & GenAI for Platform Integrity, Healthcare Diagnostics, Financial Trust and Compliance, Cybersecurity, Privacy & AI Safety: A Comprehensive Survey, Roadmap & Implementation Blueprint

Kiarash Ahi

专题命中 AI治理与伦理 :safety(title,abstract);AI safety(title);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19962 2025-07-29 cs.CL 57%

KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models

Seorin Kim, Dongyoung Lee, Jaejin Lee

机构 * Dept. of Data Science, Seoul National University(数据科学系,首尔国立大学) Dept. of Computer Science and Engineering, Seoul National University(计算机科学与工程系,首尔国立大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13138 2025-07-29 cs.CL 57%

Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation

Hadi Mohammadi, Tina Shahedi, Pablo Mosteiro, Massimo Poesio, Ayoub Bagheri, Anastasia Giachanou

机构 * Department of Methodology and Statistics, Utrecht University, The Netherlands(方法论与统计学系,乌特列支大学,荷兰) Department of Information and Computing Sciences, Utrecht University, The Netherlands(信息与计算科学系,乌特列支大学,荷兰) Queen Mary University of London, London, United Kingdom(伦敦女王玛丽大学,伦敦,英国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13175 2025-07-29 cs.AI 57%

Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era

Matthew E. Brophy

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 42 pages. Supplementary material included at end of article

详情

展开后加载摘要…

URL PDF HTML 收藏