Knowledge distillation through geometry-aware representational alignment
机构 * New York University Abu Dhabi(纽约大学阿布扎赫德分校) ; New York University(纽约大学) ; Tandon School of Engineering(Tandon工程学院)
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * New York University Abu Dhabi(纽约大学阿布扎赫德分校) ; New York University(纽约大学) ; Tandon School of Engineering(Tandon工程学院)
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI
Comments Accepted to the Cooperative Multi-Agent Systems Decision-making and Learning:Human-Multi-Agent Cognitive Fusion Workshop at AAAI 2025
机构 * Meta Superintelligence Labs(Meta 超智能实验室) ; University of Oxford(牛津大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
Comments Project page: https://junlinhan.github.io/projects/lsbs/
机构 * University of Oxford(牛津大学)
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
Comments Forthcoming in the "Reliable ML from Unreliable Data Workshop" at NeurIPS 2025
机构 * Eduard Kapelko(独立研究者)
专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG
Comments Code is available at: https://www.kaggle.com/code/kapedalex/cycleablationpublic/
机构 * Department of Data Science and its Applications, German Research Center for Artificial Intelligence (DFKI GmbH)(数据科学及其应用系,德国人工智能研究中心(DFKI GmbH)) ; Department of Computer Science, University of Kaiserslautern–Landau (RPTU)(计算机科学系,凯撒斯劳滕-兰道大学(RPTU))
专题命中 其他安全 :safety(abstract);分类 cs.AI
Comments 9 pages, 5 figures, 2 tables
机构 * University of Luxembourg(卢森堡大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
机构 * Yuqi Xiao, Yingying Zhu
专题命中 其他安全 :alignment(abstract)
专题命中 其他安全 :alignment(abstract)