arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-13 至 2025-08-13 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 8 篇

2508.05387 2025-08-13 cs.LG cs.AI 76%

Echo: Decoupling Inference and Training for Large-Scale RL Alignment on Heterogeneous Swarms

Jie Xiao, Changyuan Fan, Qingnan Ren, Alfred Long, Yuchen Zhang, Rymon Yu, Eric Yang, Lynn Ai, Shaoduo Gan

机构 * Peking University(北京大学)

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07623 2025-08-13 cs.CL 70%

Optimizing Class-Level Probability Reweighting Coefficients for Equitable Prompting Accuracy

Ruixi Lin, Yang You

机构 * Department of Computer Science(计算机科学系) National University of Singapore(新加坡国立大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08804 2025-08-13 cs.LG cs.AI 62%

TechOps: Technical Documentation Templates for the AI Act

Laura Lucaj, Alex Loosley, Hakan Jonsson, Urs Gasser, Patrick van der Smagt

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08544 2025-08-13 cs.CY cs.AI 62%

AI Agents and the Law

Mark O. Riedl, Deven R. Desai

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 2025 AAAI Conference on AI, Ethics, and Society

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05269 2025-08-13 cs.LG cs.AI q-bio.QM 62%

Chemist-aligned retrosynthesis by ensembling diverse inductive bias models

Krzysztof Maziarz, Guoqing Liu, Hubert Misztela, Austin Tripp, Junren Li, Aleksei Kornev, Piotr Gaiński, Holger Hoefling, Mike Fortunato, Rishi Gupta, Marwin Segler

机构 * Microsoft Research AI for Science(微软研究院人工智能与科学研究中心) Novartis Biomedical Research(诺华生物医学研究) University of Cambridge(剑桥大学) Jagiellonian University(雅盖隆大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08333 2025-08-13 cs.CY cs.AI 62%

Normative Moral Pluralism for AI: A Framework for Deliberation in Complex Moral Contexts

David-Doron Yaacov

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Conference version: AIES 2025 (non-archival track), 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09019 2025-08-13 cs.AI 57%

Activation Steering for Bias Mitigation: An Interpretable Approach to Safer LLMs

Shivam Dubey

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯学院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08262 2025-08-13 cs.CL 57%

Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models

Alaa Alhamzeh, Mays Al Rebdawi

机构 * First Author Affiliation(第一作者机构) Second Author Affiliation(第二作者机构)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments 8 pages, 4 figures, Passau uni, Master thesis in NLP

详情

展开后加载摘要…

URL PDF HTML 收藏