The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege
专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.CY
Comments Published at AIES 2025
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.CY
Comments Published at AIES 2025
专题命中 安全训练 :safety(abstract);分类 cs.LG
专题命中 安全训练 :safety(abstract)
Comments Joint submission paper MECC-JDSMC. Accepted for the 2025 Modeling, Estimation and Control Conference (MECC). Currently under review by the ASME Journal of Dynamic Systems, Measurement, and Control (JDSMC)
机构 * Department of Computer Science Texas A\&M University–Corpus Christi Corpus Christi, TX, USA ; EECS Department University of Missouri Columbia, MO, USA ; Department of Computer Engineering University of California–Riverside Riverside, CA, USA
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments 6 pages, 2 figures
机构 * The University of Hong Kong(香港大学)
专题命中 越狱攻击 :alignment(abstract);分类 cs.AI
Comments This paper has been accepted by the ACM Conference on Computer and Communications Security (CCS) 2025
机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) ; School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络安全学院) ; College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) ; Faculty of Data Science, City University of Macau(澳门城市大学数据科学学院)
专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI
Comments Accepted to EMNLP 2025
专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI
机构 * Singapore University of Technology and Design (SUTD)(新加坡科技设计大学) ; Institute for Infocomm Research (I2R), A*STAR, Singapore(信息通信研究院)
专题命中 安全评测 :safety(title,abstract);DPO(abstract);分类 cs.CL、cs.CY
Comments To appear at EMNLP 2025
专题命中 安全评测 :safety(title,abstract);trustworthy(abstract)
Comments 18 pages, 7 figures
专题命中 安全评测 :alignment(title,abstract)
Comments Link to publicly available codes is added
专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG
Journal ref IEEE Transactions on Software Engineering ( Volume: 49, Issue: 2, 01 February 2023)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG
机构 * Andrew Kiruluta and Priscilla Burity(独立研究者)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Peking University(北京大学) ; Microsoft(微软)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
Comments Accepted to EMNLP 2025 Main Conference
机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) ; Doxee S.p.A.(Doxee公司)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
Comments Accepted at 23rd edition of International Conference on Image Analysis and Processing 2025
机构 * Department of Electrical and Computer Engineering, University of Toronto(电气与计算机工程系,多伦多大学) ; Concordia Institute for Information Systems Engineering(康卡迪亚信息系统工程研究所) ; Concordia University(康卡迪亚大学)
专题命中 安全评测 :safety(abstract);分类 cs.AI
专题命中 安全评测 :alignment(abstract);分类 cs.LG
Comments 28 Pages, 4 Figures
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments 37 Pages,9 figures
机构 * Honda Research Institute Europe - Germany(本田欧洲研究机构)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
机构 * Department of Computer Science and Technology, Kean University, USA(计算机科学与技术系,凯恩大学,美国) ; Department of Computer Science and Technology, Wenzhou-Kean University, China(计算机科学与技术系,温州-凯恩大学,中国)
专题命中 安全评测 :trustworthy(abstract)
Comments 13 pages, 11 figures.This work has been submitted to the IEEE for possible publication
专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY
Comments 7 pages content, 1 page reference, 1 figure, Accepted at AAAI Fall Symposium Series
机构 * Principled Evolution(原则进化)
专题命中 AI治理与伦理 :alignment(abstract,comments);safety(abstract);分类 cs.CL、cs.AI、cs.CY
Comments 53 pages, 7 figures, 8 tables. Open-source implementation available at: https://github.com/Principled-Evolution/argen-demo. Work explores the integration of policy-as-code for AI alignment, with a case study in culturally-nuanced, ethical AI using Dharmic principles
机构 * Alibaba Group(阿里巴巴集团) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 其他安全 :alignment(title,abstract);safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Independent Researcher(独立研究者)
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG
Comments 5 pages, 3 figures, 1 table. tl;dr: Adversarial alignment of Time-Series Foundation Model (TSFM) embeddings enables transfer of high-quality clinical labels from medical-grade to consumer-grade wearables, enabling zero-shot prediction of gestational age without requiring paired data
机构 * AIRI ; Sber AI ; ISP RAS Research Center for Trusted AI(俄罗斯科学院信息与系统研究所可信人工智能研究中心)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments 9 pages, 1 figure, published in "The 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2025)"
Journal ref Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (2025). Association for Computing Machinery, 1152-1164
机构 * Paderborn University, Department of Business Administration and Economics(帕德博恩大学商业管理与经济学系)
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI
Comments Presented at the 19th International Conference on Wirtschaftsinformatik 2024, Würzburg, Germany https://aisel.aisnet.org/wi2024/91/
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Sri Lanka Institute of Information Technology(斯里兰卡信息技术研究所) ; Victorian Institute of Technology(维多利亚技术学院) ; Department of System Design Engineering, Keio University(系统设计工程系,庆应大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CY、cs.LG
Comments This article has been accepted for publication in IEEE Access
机构 * Student(学生) ; Downingtown STEM Academy ; Department of Computer Science(计算机科学系) ; West Chester University of Pennsylvania(宾夕法尼亚州韦斯特切斯特大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
Comments This work has been submitted to the IEEE for possible publication
专题命中 其他安全 :alignment(abstract)
Comments 27 pages, 12 figures