The Alignment Bottleneck
机构 * Independent Researcher(独立研究者)
专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Independent Researcher(独立研究者)
专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :alignment(title,abstract);分类 cs.AI
机构 * University of Minnesota, Twin Cities(明尼苏达大学,双城分校) ; University of California, Berkeley(加州大学伯克利分校)
专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI
Journal ref The 2025 Conference on Empirical Methods in Natural Language Processing
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
机构 * Universidad Diego Portales(迪亚戈·波特莱斯大学) ; Leiden University(莱顿大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments Working paper: 18 pages, 4 tables, 2 figures
Journal ref Empiria Lab Method Series (2025)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
机构 * Texas A\&M University-Corpus Christi(德克萨斯A&M大学-科珀斯克里斯蒂) ; Delft University of Technology(代尔夫特理工大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
Journal ref Proc. 2025 IEEE Conference on Artificial Intelligence (CAI), Santa Clara, CA, USA, 2025, pp. 435-440
机构 * AI Singapore(AI新加坡) ; VISTEC ; MBZUAI
专题命中 安全评测 :alignment(abstract);分类 cs.CL
Comments Accepted to EMNLP 2025 (Main). Model and Dataset: https://huggingface.co/collections/airesearch/wangchan-thai-instruction-6835722a30b98e01598984fd
专题命中 安全评测 :alignment(abstract)
Comments Accepted
Journal ref EMNLP 2025