Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants
专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY
Comments 7 pages content, 1 page reference, 1 figure, Accepted at AAAI Fall Symposium Series
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY
Comments 7 pages content, 1 page reference, 1 figure, Accepted at AAAI Fall Symposium Series
机构 * Principled Evolution(原则进化)
专题命中 AI治理与伦理 :alignment(abstract,comments);safety(abstract);分类 cs.CL、cs.AI、cs.CY
Comments 53 pages, 7 figures, 8 tables. Open-source implementation available at: https://github.com/Principled-Evolution/argen-demo. Work explores the integration of policy-as-code for AI alignment, with a case study in culturally-nuanced, ethical AI using Dharmic principles