LLM Driven Processes to Foster Explainable AI
机构 * University of Applied Sciences Ruhr West(鲁尔西部应用科学大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * University of Applied Sciences Ruhr West(鲁尔西部应用科学大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
机构 * York University(约克大学) ; Microsoft Research(微软研究院)
专题命中 其他安全 :safety(abstract);分类 cs.AI
Comments This work has been submitted to the IEEE for possible publication
专题命中 其他安全 :safety(abstract);分类 cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.AI
机构 * Algoverse AI Research(Algoverse AI研究)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
Comments Multi-Turn Interactions in Large Language Models (MTI-LLM) Workshop at NeurIPS 2025
专题命中 其他安全 :alignment(abstract);分类 cs.CL
机构 * PitchBook USA(PitchBook美国公司)
专题命中 其他安全 :alignment(abstract);分类 cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments Published in IEEE Access
Journal ref IEEE Access, vol. 13, pp. 176751-176769, 2025
机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Hong Kong University of Science and Technology(香港科学与技术大学) ; National Taiwan University, Taiwan(台湾国立台湾大学) ; Columbia University(哥伦比亚大学) ; WeBank Co., Ltd., Shenzhen, China(深圳网商银行有限公司)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; Huawei Technologies Co., Ltd.(华为技术有限公司)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments EMNLP2025 Industry Track
机构 * Institute of Computer Science, University of Tartu, Estonia(计算机科学研究所,塔尔图大学,爱沙尼亚)
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments ECAI 2025
Journal ref Frontiers in Artificial Intelligence and Applications 413 (ECAI 2025) 5027 - 5034
机构 * MPI-SWS(马克斯·普朗克所际研究所)
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments NeurIPS'25 paper
专题命中 其他安全 :alignment(abstract);分类 cs.AI
机构 * University of Texas at Arlington(德克萨斯理工大学) ; Johnson & Johnson Innovative Medicine(强生创新医学)
专题命中 其他安全 :alignment(abstract);分类 cs.LG
Journal ref NeurIPS 2025 2nd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences
机构 * Ubiquitous Knowledge Processing Lab (UKP Lab)(通用知识处理实验室) ; Department of Computer Science(计算机科学系) ; Hessian Center for AI (hessian.AI)(黑森人工智能中心)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments EMNLP 2025 Findings
机构 * NLP, Department of Computer Science, KU Leuven(自然语言处理,计算机科学系,鲁文大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments EMNLP 2025: Main Conference
机构 * Islamic University of Technology(伊斯兰技术大学)
专题命中 其他安全 :safety(abstract);分类 cs.CL
Comments Accepted at the 5th Muslims in Machine Learning (MusIML) Workshop, co-located with NeurIPS 2025
机构 * School of Engineering and Physical Sciences, Heriot-Watt University(赫瑞斯泰德大学工程与物理科学学院) ; School of Automation and Software Engineering, Shanxi University(山西大学自动化与软件工程学院) ; School of Informatics, University of Edinburgh(爱丁堡大学信息学院)
专题命中 其他安全 :safety(abstract);分类 cs.AI
机构 * Department of Industrial Engineering, Sharif University of Technology(谢里夫理工大学工业工程系) ; School of Industrial and System Engineering, Georgia Institute of Technology(佐治亚理工学院工业与系统工程学院) ; Department of Computer Science and Engineering, Ohio State University(俄亥俄州立大学计算机科学与工程系)
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments 28 pages, 18 figures
机构 * Department of Computer and Software Engineering(计算机与软件工程系)
专题命中 其他安全 :safety(abstract);分类 cs.LG
机构 * Northwestern Polytechnical University(西北工业大学) ; Nanyang Technological University(南洋理工大学) ; Chongqing University of Posts and Telecommunications(重庆邮电大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.AI
Comments Accpeted to NeurIPS 2025. Code is available at https://github.com/admins97/MSC_PRVR
机构 * Department of Linguistics University of Washington(语言学系华盛顿大学) ; Department of Linguistics Stanford University(语言学系斯坦福大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments 25 pages, 5 figures | EMNLP 2025 camera-ready version
机构 * School of Artificial Intelligence, Beijing University of Posts(人工智能学院,北京邮电大学) ; Beijing Big Data Center(北京大数据中心)
专题命中 其他安全 :safety(abstract);分类 cs.AI
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Westlake University(西交大学) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai Jiaotong University(上海交通大学) ; The Chinese University of Hong Kong(香港中文大学) ; Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
Comments 38 pages, under review
机构 * Princeton Language and Intelligence(普林斯顿语言与智能)
专题命中 其他安全 :alignment(abstract);分类 cs.LG
Comments Code available at https://princeton-pli.github.io/impossibility-unlearning/
机构 * Department of Physics, Massachusetts Institute of Technology(物理学系,麻省理工学院) ; Department of EECS, Massachusetts Institute of Technology(电子工程与计算机科学系,麻省理工学院) ; The NSF AI Institute for Artificial Intelligence and Fundamental Interactions(国家科学基金会人工智能与基本相互作用研究所)
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments NeurIPS 2025; 20 pages, 13 figures
机构 * College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments Awarded Best Student Paper at APSIPA ASC 2025
机构 * Department of Computer Science, Stanford University(斯坦福大学计算机科学系) ; Stanford Institute for Human-Centered AI(斯坦福大学人本人工智能研究所) ; MIT Media Lab(麻省理工学院媒体实验室)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)
机构 * Instituto de Computação, Universidade Estadual de Campinas (UNICAMP)(计算机学院,Campinas州立大学(UNICAMP))
专题命中 其他安全 :alignment(abstract);分类 cs.LG