arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-17 至 2025-10-17 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 隐私与版权 2 篇

2510.14312 2025-10-17 cs.AI cs.CL cs.CR 84%

Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies

Mason Nakamura, Abhinav Kumar, Saaduddin Mahmud, Sahar Abdelnabi, Shlomo Zilberstein, Eugene Bagdasarian

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) ELLIS Institute(ELLIS研究所) MPI for Intelligent Systems(智能系统研究所) Tübingen AI Center(图宾根人工智能中心)

专题命中 隐私与版权 :safety(title,abstract);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14698 2025-10-17 cs.LG cs.AI 81%

FedPPA: Progressive Parameter Alignment for Personalized Federated Learning

Maulidi Adi Prasetia, Muhamad Risqi U. Saputra, Guntur Dharma Putra

机构 * Universitas Gadjah Mada, Indonesia(加雅玛大学) Monash University, Indonesia(墨尔本大学)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 8 pages, TrustCom 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏