arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-30 至 2025-07-30 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 4 篇

2507.21091 2025-07-30 cs.CY cs.AI 81%

The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment

Lenart Motnikar, Katharina Baum, Alexander Kagan, Sarah Spiekermann-Hoff

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments Thirty-Third European Conference on Information Systems (ECIS 2025), Amman, Jordan

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21083 2025-07-30 cs.CL cs.AI 62%

ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs

Franck Bardol

机构 * Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21929 2025-07-30 cs.AI 57%

Libra: Large Chinese-based Safeguard for AI Content

Ziyang Chen, Huimu Yu, Xing Wu, Dongqin Liu, Songlin Hu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21319 2025-07-30 cs.CL 57%

Do Large Language Models Understand Morality Across Cultures?

Hadi Mohammadi, Yasmeen F. S. S. Meijer, Efthymia Papadopoulou, Ayoub Bagheri

机构 * Department of Methodology and Statistics, Utrecht University, The Netherlands(方法论与统计学系,乌特雷赫特大学,荷兰)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏