arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-10 至 2025-09-10 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 2 篇

2509.07022 2025-09-10 cs.CY cs.AI 81%

Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants

Pavan Reddy, Nithin Reddy

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

Comments 7 pages content, 1 page reference, 1 figure, Accepted at AAAI Fall Symposium Series

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07006 2025-09-10 cs.CY cs.AI cs.CL cs.LG 78%

ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code

Kapil Madan

机构 * Principled Evolution(原则进化)

专题命中 AI治理与伦理 :alignment(abstract,comments);safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 53 pages, 7 figures, 8 tables. Open-source implementation available at: https://github.com/Principled-Evolution/argen-demo. Work explores the integration of policy-as-code for AI alignment, with a case study in culturally-nuanced, ethical AI using Dharmic principles

详情

展开后加载摘要…

URL PDF HTML 收藏