arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-18 至 2025-11-18 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 5 篇

2511.12271 2025-11-18 cs.AI 85%

MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning

Zhiyu An, Wan Du

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI

Comments Accepted for AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12689 2025-11-18 cs.CY cs.AI 76%

From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM

Suyash Fulay, Jocelyn Zhu, Michiel Bakker

机构 * MIT(麻省理工学院)

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11790 2025-11-18 cs.CY cs.AI 62%

Differences in the Moral Foundations of Large Language Models

Peter Kirgis

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10089 2025-11-18 cs.LG cs.AI 62%

T2IBias: Uncovering Societal Bias Encoded in the Latent Space of Text-to-Image Generative Models

Abu Sufian, Cosimo Distante, Marco Leo, Hanan Salam

机构 * National Research Council of Italy - Institute of Applied Sciences

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments This manuscript has been accepted for presentation in the First Interdisciplinary Workshop on Responsible AI for Value Creation. Dec 1, Copenhagen. The final version will be submitted for inclusion in a Springer LNCS Volume. (The paper is 15 pages with 7 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11789 2025-11-18 cs.MA cs.AI 57%

From Single to Societal: Analyzing Persona-Induced Bias in Multi-Agent Interactions

Jiayi Li, Xiao Liu, Yansong Feng

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments AAAI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏