Co-STAR: Collaborative Curriculum Self-Training with Adaptive Regularization for Source-Free Video Domain Adaptation
机构 * University of Bristol, UK(布里斯托大学)
专题命中 幻觉与事实性 :alignment(abstract)
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * University of Bristol, UK(布里斯托大学)
专题命中 幻觉与事实性 :alignment(abstract)
专题命中 隐私与版权 :safety(abstract);AI safety(abstract);分类 cs.LG
机构 * Keio University(keio大学) ; Universitat Politècnica de València(巴塞罗那理工大学) ; Center for AI Safety(人工智能安全中心) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI
机构 * Rochester Institute of Technology(罗切斯特理工学院) ; U.S. Naval Research Laboratory(美国海军研究实验室) ; University of Missouri-Kansas City(密苏里大学-Kansas城分校) ; Adobe(Adobe公司) ; Meituan(美团) ; Sun Yat-sen University(孙中山大学) ; Purdue University(普渡大学) ; UC Davis(加州大学戴维斯分校) ; University of Rochester(罗切斯特大学) ; Rice University(莱斯大学) ; Meta AI
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments NeurIPS 2025
专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG
Comments Accepted for publication in Frontiers in Artificial Intelligence - Medicine and Public Health (Original Research)
Journal ref Front. Artif. Intell. 8:1662984 (2025)
机构 * University of Southern California(南加州大学) ; USC for Institute of Creative Technologies(南加州大学创意技术研究所)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments Accepted to EMNLP 2025 (Main Conference)
机构 * The Chinese University of Hong Kong(香港中文大学) ; City University of Hong Kong(香港城市大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室) ; University of Edinburgh(爱丁堡大学)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL
Comments This work has been accepted by EMNLP 2025
机构 * Amazon Nova Responsible AI(亚马逊Nova负责任的人工智能) ; Center for AI Safety(人工智能安全中心) ; CMU(卡内基梅隆大学) ; Gray Swan AI(灰天鹅人工智能)
专题命中 安全评测 :alignment(abstract);safety(abstract);prompt injection(abstract);分类 cs.CL
Comments Preprint
机构 * The Graduate University for Advanced Studies (SOKENDAI)(高级研究大学(SOKENDAI)) ; National Institute of Informatics(信息研究所)
专题命中 安全评测 :alignment(title);分类 cs.CL
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)
专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI
Comments 34 pages in total, EMNLP 2025
机构 * Graduate School of Information Science, University of Hyogo(京都大学垣田学园信息科学研究生院) ; Center for Informatics Science, School of Information Technology and Computer Science, Nile University(尼罗大学信息科学中心)
专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI
Comments NeurIPS2025 Workshop: Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling
机构 * University of Maryland, College Park(马里兰大学科帕克分校) ; City University of Hong Kong(香港城市大学) ; University of Southern California(南加州大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
Comments NeurIPS 2025
机构 * Nankai University(南开大学) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai AI Laboratory(上海人工智能实验室) ; Fudan University(复旦大学) ; Johns Hopkins University(约翰霍普金斯大学) ; Wuhan University(武汉大学) ; University of Science and Technology of China(中国科学技术大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments EMNLP 2025 MainConference
机构 * University of the Peloponnese, Department of Informatics and Telecommunications(希腊皮埃蒙特大学信息与电信系) ; Cornell University, Department of Biological and Environmental Engineering(康奈尔大学生物与环境工程系)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
机构 * Center for Intelligent Information Retrieval(智能信息检索中心) ; University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG
机构 * Northeastern University(东北大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments This is a preprint version of the paper accepted to IVA'25
机构 * University of Maryland(马里兰大学)
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG
Comments Project website: http://vulnerable-ai-agents.github.io
机构 * Worcester Polytechnic Institute(沃斯特理工大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
机构 * National Yang Ming Chiao Tung University(国家阳明交通大学) ; NVIDIA
专题命中 安全评测 :alignment(abstract);分类 cs.CL
Comments Accepted for EMNLP 2025 main conference
机构 * NUS(国立新加坡大学) ; SonarSource SA
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments 4 pages
机构 * University of Rostock(罗斯托克大学) ; University of Rostock, University of Marburg(罗斯托克大学、马堡大学)
专题命中 安全评测 :alignment(abstract);分类 cs.LG
Comments 10 pages, 7 figures, added repo link
机构 * EPFL(苏黎世联邦理工学院) ; MIT(麻省理工学院) ; Georgia Institute of Technology(佐治亚理工学院)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
Comments EMNLP 2025. Project Page at https://language-to-cognition.epfl.ch
机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(莫扎伊德大学人工智能学院) ; University of Notre Dame(诺特丹大学) ; IBM Research(IBM研究院)
专题命中 安全评测 :DPO(abstract);分类 cs.CL
机构 * Department of Computer Science and Engineering, University of Notre Dame(诺丁汉大学计算机科学与工程系) ; MIT(麻省理工学院) ; CalTech(加州理工学院) ; MBZUAI(穆斯林人工智能研究所) ; CMU(卡内基梅隆大学) ; Department of Chemistry & Biochemistry, University of Notre Dame(诺丁汉大学化学与生物化学系)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
机构 * Computer Science Department(计算机科学系) ; Applied Economics Department(应用经济学系)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments 6 pages, 3 figures, 4 tables, 1 algorithm, accepted in the Robustness and Security of Large Language Models (ROSE-LLM) special session at ICMLA 2025
专题命中 安全评测 :safety(abstract);分类 cs.AI
机构 * Laboratory for Emerging Intelligence(新兴智能实验室) ; University of California, San Diego(加州大学圣地亚哥分校)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
Comments Published at EMNLP 2025 (Main)
专题命中 安全评测 :trustworthy(abstract)
Comments 17 pages, 7 figures
机构 * College of Computing(计算学院) ; Georgia Institute of Technology(佐治亚理工学院) ; University of Illinois(伊利诺伊大学) ; Neuro Industry Research(神经产业研究) ; Neuro Industry, Inc.(神经产业公司)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG
机构 * institutetext(机构文本) ; CNES(法国国家空间研究中心) ; INSA-IMT(法国里尔INSA-IMT)
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
Comments Accepted as a workshop paper at MACLEAN - ECML/PKDD 2025