Back to the Future -- Sequential Alignment of Text Representations
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL
Comments AAAI 2020
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL
Comments AAAI 2020
专题命中 其他安全 :safety(title,abstract);分类 cs.LG
Comments Revised text (2nd version). The research work is conducted during the first author's service at the fortiss research institute and is supported by the following projects: "Audi Verifiable AI" from Audi AG, Germany and "Dependable AI for automotive systems" from DENSO Corporation, Japan
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
专题命中 其他安全 :safety(title,abstract);分类 cs.LG
Comments Tool available at https://github.com/dependable-ai/nn-dependability-kit
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
Comments This paper has been accepted by AAAI-2019
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
Comments NIPS workshop on Multi-Modal Machine Learning, 2015
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
Comments Joint European Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD, 2014
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL
专题命中 其他安全 :safety(title,abstract);分类 cs.AI
Comments Extended Abstract, 3 pages, Accepted at IBM Collaborative Academia Research Exchange (I-CARE)-2011, uses ACM-Proceeding style file
专题命中 其他安全 :safety(title,abstract);分类 cs.AI
Journal ref Journal Of Artificial Intelligence Research, Volume 17, pages 363-378, 2002
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI
Journal ref Proceedings of the Workshop and Tutorial on Learning Context-Free Grammars (in association with the 14th European Conference on Machine Learning and the 7th European Conference on Principles and Practice of Knowledge Discovery in Databases (ECML/PKDD 2003), September 2003, Cavtat-Dubrovnik, Croata), editors: C. de la Higuera and P. Adriaans and M. van Zaanen and J. Oncina, pp 113-124
作为人工智能-智能体对齐层次的基础市场设计
专题命中 其他安全 :alignment(title,abstract)
AI总结 研究探讨市场中人工智能-智能体对齐,提出将基础市场设计视为该对齐层次,通过对市场核心形式化建模,利用理论计算机科学严谨性构建透明盒模型,支持激励分析与机制设计,使期望行为受青睐,不良行为难维持。
Comments Accepted as at the EC'26 Workshop on Incentive-Based AI Alignment, co-located with the 27th ACM Conference on Economics and Computation, Rome, Italy, July 2026. This version is prepared for public dissemination following workshop acceptance
视觉语言模型的深度预对齐
机构 * Tsinghua University ; Shanghai Qi Zhi Institute ; Taobao \& Tmall Group of Alibaba
专题命中 其他安全 :alignment(title,abstract)
AI总结 本文提出深度预对齐(DPA),通过替换传统ViT编码器为小型VLM作为感知器,实现视觉特征与目标大语言模型文本空间的深度对齐,提升了多模态基准性能,并降低了语言能力遗忘。
Comments Accepted by ICML 2026. Project Website: https://github.com/THUMAI-Lab/Deep-Pre-Alignment
专题命中 其他安全 :alignment(title,abstract)
Comments GitHub: https://github.com/SHI-Labs/Diffusion-Driven-Test-Time-Adaptation-via-Synthetic-Domain-Alignment
专题命中 其他安全 :safety(title,abstract)
Journal ref Safety Science, 2024, 177, pp.106576
专题命中 其他安全 :alignment(title,abstract)
Comments Update Tri-modal Alignment task
评估大型语言模型在复杂隐藏角色游戏中的表现
机构 * University of Göttingen(哥廷根大学)
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI
AI总结 本研究通过社交推理游戏《秘密希特勒》评估大型语言模型的推理、说服和欺骗能力,引入新指标并发现当前模型在复杂多轮操纵中效果不佳。
Comments Master's thesis, University of Göttingen
多智能体AI中的隐藏联盟:来自内部表示的谱诊断
机构 * Reciprocal Research(递归研究) ; Center for the Future of AI, Mind, and Society(人工智能、心智与社会未来中心) ; Florida Atlantic University(佛罗里达 Atlantic 大学) ; Biological and Computational Intelligence Center(生物与计算智能中心) ; National Intelligence University(国家情报大学)
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG
AI总结 本文提出通过分析多智能体系统内部神经表示的谱分区方法,检测隐藏联盟结构,验证了该方法在强化学习和大语言模型中的有效性,揭示了代表层次结构。
Comments 18 pages
信息论视角下欺骗与混淆的区别
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG
AI总结 本文从信息论角度区分两种AI安全失效模式:欺骗对齐与目标漂移,揭示二者在人类-AI系统不同接口的信息分歧,提出形式化模型和思想实验,为大型语言模型对齐挑战提供新视角。
Comments Proceedings of the 14th IJCNLP and the 4th AACL (2025)
语言模型中的地位层级
机构 * Brigham Young University–Hawaii ; COLUMBIA UNIVERSITY(哥伦比亚大学)
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI
AI总结 语言模型在多智能体环境中会因地位线索形成层级,高地位分配反而降低高能力模型的服从,揭示AI系统中的新兴社会行为。
机构 * Department of Philosophy University of North Carolina at Chapel Hill(哲学系北卡罗来纳大学教堂山分校) ; Department of Computer Science University of North Carolina at Chapel Hill(计算机科学系北卡罗来纳大学教堂山分校) ; Department of Computer Science University of Texas at Austin(计算机科学系德克萨斯大学奥斯汀分校)
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI
Comments substantial expansions of sections 4 and 5, updated references, numerous smaller additions and clarifications
机构 * Stanford University(斯坦福大学) ; Mathematical Medicine Group(数学医学组) ; Department of Neurosurgery(神经外科系) ; Physician-Scientist Training Program(医师科学家培训计划)
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.CY
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG
面向关键任务中故障可运行无人机集群的安全驱动架构框架
专题命中 其他安全 :safety(title,abstract)
AI总结 针对安全关键操作中无人机集群的认证需求,提出结合SAE ARP4754B方法的混合关键度架构框架,通过硬件隔离的安全监控器实现飞行关键核心与集群管理器解耦,经马尔可夫建模可满足危险故障要求。
Comments 10 pages, 7 figures. Accepted for presentation at the 45th AIAA/IEEE Digital Avionics Systems Conference (DASC), Orlando, FL, USA, 2026. \c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses
ConceptFormer:学习自适应潜在概念以实现视觉文档检索中的查询-文档对齐
机构 * Northeastern University(东北大学) ; Tsinghua University(清华大学) ; Peking University(北京大学)
专题命中 其他安全 :alignment(title,abstract)
AI总结 针对视觉文档检索中现有监督信号的局限,本文提出ConceptFormer框架,以自适应潜在概念为中间表示衔接语义鸿沟,在基准测试中较最强基线实现了16.7%、22.1%的NDCG@10相对提升,性能优异。
深度模型,浅层对齐:揭示神经解码中的粒度不匹配
机构 * Dept. of Electrical & Computer Engineering, University of Pittsburgh, USA(宾夕法尼亚大学电气与计算机工程系) ; Dept. of Biomedical Engineering, Tsinghua University, China(清华大学生物医学工程系) ; Dept. of Neurology, University of Southern California, USA(美国南加州大学神经病学系) ; Dept. of Computer Science, University of Texas Rio Grande Valley, USA(德克萨斯理工大学里奥格兰德谷分校计算机科学系)
专题命中 其他安全 :alignment(title,abstract)
AI总结 本文提出浅层对齐方法,通过对比学习策略解决神经解码中的粒度不匹配问题,显著提升解码性能。
Comments 33 pages, 16 figures
Journal ref Transactions on Machine Learning Research (2026), ISSN 2835-8856
基于生成式障碍证书的非线性系统形式化安全验证
专题命中 其他安全 :safety(title,abstract)
AI总结 该研究针对非线性系统安全验证中障碍证书推导计算成本高的问题,提出基于大语言模型的生成式框架,将BMI问题转化为LMI测试,实现了远超传统方法的速度与性能提升。