Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Tencent Inc.(腾讯公司)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG
机构 * Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(合作中位网创新中心,上海交通大学) ; School of Artificial Intelligence, Shanghai Jiao Tong University(人工智能学院,上海交通大学) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Beijing Institute of Technology(北京理工大学) ; A*STAR Centre for Frontier AI Research(A*STAR前沿人工智能研究中心)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
Comments 16 pages, Accepted at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
专题命中 其他安全 :alignment(abstract)
Comments 9 pages, 7 figures, 4 tables
专题命中 其他安全 :safety(abstract)
专题命中 其他安全 :alignment(abstract)
Comments 19 pages, 6 figures