Building Trust: Foundations of Security, Safety and Transparency in AI
专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted by NeurIPS 2024
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments This work has been submitted to the ELSEVIER for possible publication
专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
Journal ref 43rd Annual CNS Conference and the 48th Annual CNS/CNA Student Conference Sheraton Cavalier Saskatoon Hotel, Saskatoon, SK, Canada, June 16-19, 2024
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments RecSys 2024 (Long Paper)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to Findings of the Association for Computational Linguistics: ACL 2024
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments CogSci
专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 34 pages, 7 figures, 2 tables
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 7 pages, 5 figures
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments To appear in Proceedings of the 1st Personalization of Generative AI Workshop, EACL 2024
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY、cs.LG
Comments 127 pages, 8 figures. Revised again to correct typos
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 22 pages, 19 Figures, 7 Tables
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted at ISWC 2023 (Posters and Demos)
专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Submitted to Nature - Machine Intelligence (Revised and Extended)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments LREC 2022 Main Conference Accepted Paper
专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY、cs.LG
Comments 3 pages
基于自监督视觉Transformer与协同跨域对齐的高效无监督域适应
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG
AI总结 提出EUDA框架,以冻结的DINOv2为特征提取器,结合SDAL损失,在多数据集上实现高效无监督域适应,可训练参数减少42%至99.7%,适配资源受限环境。
Comments 22 pages, 4 figures
Journal ref Abedi, A., Wu, Q.M.J., Zhang, N. et al. Efficient unsupervised domain adaptation via self-supervised vision transformer and synergistic cross-domain alignment. Int. J. Mach. Learn. & Cyber. 17, 423 (2026)
机构 * Independent Researcher(独立研究者)
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG
Comments 5 pages, 3 figures, 1 table. tl;dr: Adversarial alignment of Time-Series Foundation Model (TSFM) embeddings enables transfer of high-quality clinical labels from medical-grade to consumer-grade wearables, enabling zero-shot prediction of gestational age without requiring paired data
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY
Comments 7 pages, 8 figures, 3 tables, forthcoming at the AAAI-25 Special Track on AI Alignment
专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.CY
Comments Available under the open government license at https://www.gov.uk/government/publications/international-scientific-report-on-the-safety-of-advanced-ai
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments Paper List: https://github.com/cascip/awesome-auto-alignment
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments Code and data: https://github.com/joyheyueya/psychometric-alignment
专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG
Comments International Conference on Computer Safety, Reliability, and Security 2023
专题命中 其他安全 :safety(title,abstract);分类 cs.CY、cs.LG
Comments DISE1: Joint Workshop on Deep (or Machine) Learning for Safety-Critical Applications in Engineering
TextLDM:基于连续潜在扩散的语言建模
机构 * Joy Future Academy(京东探索研究院) ; HIT(Harbin Institute of Technology) ; HKUST(GZ)(Hong Kong University of Science and Technology (Guangzhou))
专题命中 其他安全 :alignment(summary_cn,abstract);分类 cs.CL
AI总结 TextLDM将视觉潜在扩散框架应用于文本生成,通过Representation Alignment提升文本表示质量,在OpenWebText2上训练后优于现有扩散语言模型,匹配GPT-2性能。
M-DaQ:用于指令微调数据集的多语言多样性与质量样本检索
机构 * Huawei Technologies Ltd.(华为技术有限公司) ; University of Science and Technology of China(中国科学技术大学)
专题命中 其他安全 :alignment(summary_cn,abstract);分类 cs.CL
AI总结 M-DaQ通过联合优化指令-响应质量与跨语言语义多样性,构建高质量平衡训练数据,验证了多语言设置下的Superficial Alignment Hypothesis,并在18种语言上展示出超过60%的胜率。
Comments Accepted by SIGIR 2026 Short
面向参数化安全规范的高效动态防护机制
机构 * University of California, Irvine(加州大学尔湾分校) ; IMDEA Software Institute(IMDEA软件研究所)
专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG
AI总结 本文针对参数化安全规范提出动态防护机制,其离线设计耗时数分钟,在线适配速度较暴力重新计算方法快最多5倍,可用于未知区域机器人导航以应对安全规范的动态变化。
Journal ref International Symposium on Automated Technology for Verification and Analysis (ATVA) 2025, pp. 157-179. Cham: Springer Nature Switzerland