arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-07 至 2025-11-07 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 3 篇

2511.04157 2025-11-07 cs.SE cs.AI 83%

Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment

Asma Yamani, Malak Baslyman, Moataz Ahmed

机构 * Information and Computer Science Department, KFUPM(信息与计算机科学系,KFUPM) IRC for finance and digital economy, KFUPM(金融与数字经济研究中心,KFUPM) SDAIA-KFUPM Joint Research Center for Artificial Intelligence, KFUPM(人工智能联合研究中心,KFUPM)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02895 2025-11-07 cs.CY cs.AI cs.HC physics.soc-ph 73%

A Criminology of Machines

Gian Maria Campedelli

机构 * Fondazione Bruno Kessler(布鲁诺·科斯勒基金会)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments This pre-print is also available at CrimRxiv with DOI: https://doi.org/10.21428/cb6ab371.e3354ce1

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03980 2025-11-07 cs.AI cs.CL 62%

LLMs and Cultural Values: the Impact of Prompt Language and Explicit Cultural Framing

Bram Bulté, Ayla Rigouts Terryn

机构 * Brussels Centre for Language Studies, Vrije Universiteit Brussel(布鲁塞尔语言研究中心,布鲁塞尔自由大学) Université de Montréal & Mila - Quebec AI Institute(蒙特利尔大学及魁北克人工智能研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Preprint under review at Computational Linguistics. Accepted with minor revisions (10/10/2025); second round

详情

展开后加载摘要…

URL PDF HTML 收藏