Law Informs Code: A Legal Informatics Approach to Aligning Artificial Intelligence with Humans
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Northwestern Journal of Technology and Intellectual Property, Volume 20, Issue 3, 2023
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Northwestern Journal of Technology and Intellectual Property, Volume 20, Issue 3, 2023
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Code and models are available at https://github.com/ZrrSkywalker/LLaMA-Adapter
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Published at NeurIPS 2022
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments 30 pages, 10 figures, 5 tables. Website: https://sites.google.com/view/robots-enact-stereotypes . Published in the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT 22), June 21-24, 2022, Seoul, Republic of Korea. ACM, DOI: https://doi.org/10.1145/3531146.3533138 . FAccT22 Submission dates: Abstract Dec 13, 2021; Submitted Jan 22, 2022; Accepted Apr 7, 2022
Journal ref In 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT 22). ACM, New York, NY, USA, 743-756
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Published in FAccT 2022
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
Comments 39 Pages
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
Comments AISTATS 2022. Code is released at https://github.com/JFChi/Return-Parity-MDP
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY
Comments published in Journal of Artificial Intelligence Research, 71: 431-478, July 2021
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (AIES 2021)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
对 Machiavellian 代理进行对齐:通过测试时策略塑造实现行为引导
专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CL、cs.AI
AI总结 本文提出了一种测试时策略塑造方法,通过模型引导的策略调整,解决预训练代理在复杂环境中的伦理对齐问题,实现奖励最大化与伦理约束的平衡。
Comments Accepted to AAAI 2026 AI Alignment Track
机构 * Leipzig University(莱比锡大学) ; Technical University Dresden(德累斯顿技术大学) ; University of Göttingen(哥廷根大学)
专题命中 AI治理与伦理 :trustworthy(abstract,comments);分类 cs.AI、cs.CY
Comments Accepted paper - ESORICS 2025 - International Workshop on Secure and Trustworthy Machine Unlearning Systems (STMUS)
机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Shenzhen Research Institute of Big Data(深圳大数据研究院)
专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CL、cs.AI
Comments Cultural Analysis, Cultural Alignment, Word Association Test, Large Language Models. Accepted by EMNLP 2025 (Oral)
机构 * Tier University(Tier大学)
专题命中 AI治理与伦理 :alignment(abstract,journal_ref);分类 cs.CL、cs.AI
Comments 15pages, 1 figure, 2 tables
Journal ref Proceedings of 0th Symposium on Moral and Legal AI Alignment of the IACAP/AISB Conference, 2025
专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.AI、cs.LG
Comments AAAI-25 AI Alignment Track
专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.AI、cs.LG
Comments The current version improves significantly by integrating ethical frameworks, expanding methodology and case studies, enhancing scalability and ethical-legal alignment, acknowledging prior work, and offering clearer structure and practical relevance
专题命中 AI治理与伦理 :trustworthy(abstract,comments);分类 cs.AI、cs.CY
Comments Accepted at 'Neural Conversational AI Workshop - What's left to TEACH (Trustworthy, Enhanced, Adaptable, Capable, and Human-centric) chatbots?' at ICML 2023
时间任务多样性:非平稳性下的归纳偏置
机构 * University of Oxford(牛津大学)
专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.LG;AI safety(comments)
AI总结 研究探讨了在合成序列建模中,任务分布随时间变化对深度学习模型归纳偏置的影响,发现任务分布的多样性增强了模型对泛化而非记忆的偏好。
Comments Presented at Technical AI Safety Conference (TAIS), Oxford, May 2026. Code available at https://github.com/matomatical/temporal-task-diversity
锚定偏差:一种针对持续学习下多模态大语言模型(MLLMs)的持久公平性后门攻击
机构 * Emory University(埃默里大学)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG
AI总结 针对持续学习下的多模态大语言模型,研究人员提出持久公平性后门攻击,通过两种机制注入持久群体歧视,该攻击能规避标准防御且在多轮持续学习中留存。
Comments CIKM 2026
MAVEN-T:用于实时多智能体轨迹预测的强化异构蒸馏
机构 * School of Mathematical Sciences, Shanghai Jiao Tong University(上海交通大学数学科学学院) ; Bio-X Institutes, Key Laboratory for the Genetics of Developmental and Neuropsychiatric Disorders, Shanghai Jiao Tong University(上海交通大学Bio-X研究院、发育与神经精神疾病遗传学重点实验室) ; Shanghai Key Laboratory of Psychotic Disorders, Brain Science and Technology Research Center, Shanghai Jiao Tong University(上海精神疾病重点实验室、脑科学与技术研究中心,上海交通大学)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG
AI总结 提出MAVEN-T框架,通过高容量教师模型和紧凑学生模型的异构蒸馏,结合强化学习优化,实现实时多智能体轨迹预测,在多个数据集上达到高精度与低延迟。
大型语言模型的六大误解:一个极简模型与诊断分类法
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
AI总结 该研究提出以四组区分为核心的LLM极简模型,诊断六大误解,应用于出版商AI政策案例,为纠正民间理论错误提供诊断工具。
Comments 20 pages, 1 figure, 2 tables, and 2 boxes. Published in PNAS Nexus
Journal ref PNAS Nexus, 5(7), pgag236 (2026)
2026年全球负责任AI指数:概念框架与方法论
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY
AI总结 该研究介绍2026年全球负责任AI指数(GIRAI)第二版的方法论,优化框架维度与指标,经审计验证,用于跨国评估各国负责任AI治理,助力相关主体识别保护成效与差距。
在安全调查中使用和评估大语言模型(LLM)作为代理专家的框架:可靠性、偏差及启示
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
AI总结 本研究提出了评估LLM作为安全调查代理专家的框架,发现LLM虽内部一致但与专家响应存在系统性偏差,可用于试点和假设生成但不能替代专家征询。
智能体AI的运行时治理:基于可信溯源与故障闭锁执行的行动边界控制
机构 * SPQR Technologies Inc.(SPQR科技公司)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY
AI总结 该研究提出Aegis运行时治理系统,通过可信决策层调解智能体AI的工具行动提案,在沙堡语料库评估中成功阻止风险提案转化为治理副作用。
ETHOS:面向临床多智能体系统的模块化伦理框架
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG
AI总结 ETHOS是可与现有临床多智能体系统集成的模块化伦理框架,通过分层治理提升决策可靠性,将AI伦理原则转化为可部署的安全保障。
Comments Preprint of an article submitted for consideration in Pacific Symposium on Biocomputing \textcopyright\ 2027 World Scientific Publishing Company. \url{https://psb.stanford.edu/}
文明超材料:能力梯度与结构湍流下的协调工程
机构 * Independent Researcher(独立研究者)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
AI总结 受超材料物理学启发,提出将治理从规范性学科转变为工程学科的正式框架,通过有效协调系数模型预测自愈与自失稳相变,并设计可检验假设与实验方案。
Comments 19 pages, 4 figures. Accepted for presentation at AGI-26 (Springer LNAI, forthcoming). v2 corrects the sign of the synergy term in the constitutive law (Eq. 2) and reformulates H3 as a threshold-crossing claim, per peer review
Journal ref Artificial General Intelligence. AGI 2026. Lecture Notes in Computer Science, vol 16855, pp. 118-136. Springer, Cham (2026)
虚假朋友困境:信任与对话AI的政治经济学
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
AI总结 本文提出虚假朋友困境框架,探讨拟人化AI如何通过隐蔽手段影响用户自主权,分析其在信任与政治经济学中的作用。
Comments Manuscript under review
协同构建社会技术型AI治理:利用算法登记册的参与式系统映射
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY
AI总结 本文以荷兰某城市算法登记册为案例,通过多利益相关者参与式映射结合STPA分析,揭示登记册的遮蔽性,为多元社会技术型AI治理提供新视角。