MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning
专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI
Comments Accepted for AAAI 2026
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI
Comments Accepted for AAAI 2026
机构 * MIT(麻省理工学院)
专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.CY
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
机构 * National Research Council of Italy - Institute of Applied Sciences
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
Comments This manuscript has been accepted for presentation in the First Interdisciplinary Workshop on Responsible AI for Value Creation. Dec 1, Copenhagen. The final version will be submitted for inclusion in a Springer LNCS Volume. (The paper is 15 pages with 7 figures)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI
Comments AAAI-2026