Evaluating Human-LLM Representation Alignment: A Case Study on Affective Sentence Generation for Augmentative and Alternative Communication
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL
Comments Published at IJCNLP-AACL 2025 Findings
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL
Comments Published at IJCNLP-AACL 2025 Findings
专题命中 安全评测 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CY
Comments Accepted to the Workshop on Regulatable ML at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)
机构 * University of California, Berkeley(加州大学伯克利分校) ; University of California, Santa Cruz(加州大学圣克ruz分校) ; University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; University of Chicago(芝加哥大学) ; Amazon(亚马逊) ; Information Commissioner’s Office(信息专员办公室)
专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.LG
Comments NeurIPS 2025 Datasets & Benchmarks
机构 * Independent(独立研究者) ; UC Irvine(加州大学尔湾分校) ; University of Waterloo(滑铁卢大学) ; UW Madison(威斯康星大学麦迪逊分校) ; University of Toronto(多伦多大学) ; Algoverse(Algoverse公司) ; META ; University of Maryland(马里兰大学)
专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI
机构 * NLP Team, Xiaohongshu Inc.(小红书研究院自然语言处理团队)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Innov-Acts Ltd.(Innov-Acts有限公司) ; CYENS Centre of Excellence(CYENS卓越中心) ; University of Piraeus(比雷埃克斯大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 7 pages, 7 figures
Journal ref 2025 21st International Conference on Distributed Computing in Smart Systems and the Internet of Things (DCOSS-IoT)
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
Comments McGill PhD Thesis (updated on 20251109 for typos and margin adjustments)
机构 * University of Texas at Austin, TX, USA(德克萨斯大学奥斯汀分校)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 8 PAGES
机构 * Inclusion AI Ant Group(Inclusion AI Ant集团)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 32 pages, 8 figures
机构 * Warsaw University of Technology(华沙技术大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
Comments Accepted for publication at the 34th ACM International Conference on Information and Knowledge Management (CIKM '25)
Journal ref CIKM '25: Proceedings of the 34th ACM International Conference on Information and Knowledge Management (2025) 2535-2545
机构 * Department of Electronics and Communication Engineering, Istanbul Technical University(电子与通信工程系,伊斯坦布尔技术大学)
专题命中 安全评测 :safety(abstract);分类 cs.AI
专题命中 安全评测 :alignment(abstract);分类 cs.LG
机构 * University of Notre Dame(诺特大学) ; Vanderbilt University(范德比大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
机构 * University of North Texas(北卡罗来纳州立大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
Comments Accepted for publication at 2nd ACM International Conference on AI-powered Software (AIware 2025)
机构 * King Saud University, College of Computer and Information Sciences(沙特王后大学,计算机与信息科学学院)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
Comments Accepted, to appear IJCNLP-AACL 2025 Findings
机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) ; Carnegie Mellon University(卡内基梅隆大学) ; DatologyAI
专题命中 安全评测 :safety(abstract);分类 cs.CL
专题命中 安全评测 :safety(abstract);分类 cs.AI
Comments AAAI 2026 (Senior Track)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
机构 * SKLCCSE, School of Computer Science and Engineering, Beihang University(北京航空航天大学信息与电子技术学院) ; Key Lab of Education Blockchain and Intelligent Technology, Guangxi Normal University(广西师范大学教育区块链与智能技术重点实验室) ; School of Computing, National University of Singapore(新加坡国立大学计算机学院) ; Department of Computer Science, University of Illinois, Chicago(伊利诺伊大学芝加哥分校计算机科学系)
专题命中 安全评测 :trustworthy(abstract);分类 cs.LG
Comments Accepted by the NeurIPS 2025
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments 24pages
机构 * UC Santa Barbara(加州大学圣芭芭拉分校) ; NIKSUN, Inc(NIKSUN公司)
专题命中 安全评测 :alignment(abstract);分类 cs.LG
专题命中 安全评测 :safety(abstract);分类 cs.AI
Comments in Chinese language
机构 * BRIA AI(BRIA人工智能)
专题命中 安全评测 :alignment(abstract)
机构 * Department of Artificial Intelligence Convergence(人工智能融合系)
专题命中 安全评测 :alignment(abstract)
专题命中 安全评测 :trustworthy(abstract)
Comments 10 pages, 5 figures, and 5 tables
机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK–Shenzhen)(数据科学学院,香港中文大学(深圳)) ; Computer Science Program, CEMSE Division, King Abdullah University of Science and Technology (KAUST)(计算机科学项目,科学与工程学院,国王 Abdullah 科学技术大学) ; Center of Excellence on Smart Health, KAUST(智能健康卓越中心,国王 Abdullah 科学技术大学) ; Center of Excellence for Generative AI, KAUST(生成式人工智能卓越中心,国王 Abdullah 科学技术大学) ; Department of Dermatology, Beijing AnZhen Hospital, Capital Medical University(皮肤科,北京安贞医院,首都医科大学) ; Department of Dermatology, Tianjin Institute of Integrative Dermatology, Tianjin Academy of Traditional Chinese Medicine Affiliated Hospital(皮肤科,天津整合皮肤科研究院,天津中医药大学附属医院) ; Department of Dermatology, Beijing Aerospace General Hospital(皮肤科,北京航天总医院) ; Department of Dermatology, The First Affiliated Hospital, Shantou University Medical College(皮肤科,汕头大学医学院第一附属医院) ; DermAssure, LLC(DermAssure 公司) ; School of Medicine, New York Medical College(医学院,纽约医学院) ; Capital Medical University(首都医科大学) ; Department of Dermatology, Second Hospital of Jilin University(皮肤科,吉林大学第二医院) ; Emergency Critical Care Center, Beijing AnZhen Hospital, Capital Medical University(急诊重症中心,北京安贞医院,首都医科大学)
专题命中 安全评测 :trustworthy(abstract)
机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)科学与工程学院) ; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院) ; Shenzhen Institute of Artificial Intelligence and Robotics for Society, China(深圳人工智能与机器人研究院)
专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL
机构 * Marquette University, WI, USA(马quette大学)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI
Comments Accepted in IEEE Big Data, 8-11 December, 2025 @ Macau SAR, China
专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.CL、cs.LG
Comments refactor