arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-24 至 2025-10-24 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 3 篇

2505.09131 2025-10-24 cs.LG cs.AI stat.ML 81%

Fair Clustering via Alignment

Kunwoong Kim, Jihu Lee, Sangchul Park, Yongdai Kim

机构 * Department of Statistics, Seoul National University, Republic of Korea(统计系,首尔国立大学,大韩民国) School of Law, Seoul National University, Republic of Korea(法学院,首尔国立大学,大韩民国)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

Journal ref ICML 2025 (Forty-Second International Conference on Machine Learning)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20782 2025-10-24 cs.CL cs.AI 62%

A Use-Case Specific Dataset for Measuring Dimensions of Responsible Performance in LLM-generated Text

Alicia Sagae, Chia-Jung Lee, Sandeep Avula, Brandon Dang, Vanessa Murdock

机构 * AWS Responsible AI(AWS负责任人工智能)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments 24 pages with 3 figures, to appear in Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19733 2025-10-24 cs.CL cs.LG 62%

Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning

M. H. I. Abdalla, Zhipin Wang, Christian Frey, Steffen Eger, Josif Grabocka

机构 * Department of Computer Science University of Technology Nuremberg(计算机科学系图腾技术大学纽伦堡)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏