Multi-Prompt Progressive Alignment for Multi-Source Unsupervised Domain Adaptation
机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身人工智能研究院,复旦大学)
专题命中 AI治理与伦理 :alignment(title,abstract)
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身人工智能研究院,复旦大学)
专题命中 AI治理与伦理 :alignment(title,abstract)
机构 * School of Software Engineering, South China University of Technology(南方科技大学软件工程学院) ; South China University of Technology(南方科技大学) ; Pazhou Lab(琶洲实验室) ; School of Future Technology, South China University of Technology(未来技术学院) ; Peng Cheng Laboratory(鹏城实验室) ; Key Laboratory of Big Data and Intelligent Robot (South China University of Technology), Ministry of Education(大数据与智能机器人重点实验室)
专题命中 AI治理与伦理 :alignment(title,abstract)
Comments This paper is accepted by IEEE TIP 2025 (The journal version is available at https://doi.org/10.1109/TIP.2025.3586487). Code is publicly available at https://github.com/kaai520/PGFA
Journal ref IEEE Transactions on Image Processing 34 (2025) 4602-4617
专题命中 AI治理与伦理 :alignment(title,abstract)
Journal ref ACM Conference on Fairness, Accountability, and Transparency 2025 (ACM FAccT 2025)
专题命中 AI治理与伦理 :safety(title,abstract)
专题命中 AI治理与伦理 :alignment(title,abstract)
专题命中 AI治理与伦理 :safety(title,abstract)
专题命中 AI治理与伦理 :alignment(title,abstract)
专题命中 AI治理与伦理 :alignment(title,abstract)
Comments Project page: https://ai.stanford.edu/~yzzhang/projects/3d-congealing/
专题命中 AI治理与伦理 :trustworthy(title);分类 cs.AI、cs.CY、cs.LG
专题命中 AI治理与伦理 :alignment(title,abstract)
Comments 20 pages, 23 figures
专题命中 AI治理与伦理 :trustworthy(title);分类 cs.AI、cs.CY、cs.LG
专题命中 AI治理与伦理 :trustworthy(title);分类 cs.AI、cs.CY、cs.LG
Comments 46 pages
专题命中 AI治理与伦理 :safety(title);分类 cs.AI、cs.CY、cs.LG
Journal ref The Alan Turing Institute (June, 2019)
专题命中 AI治理与伦理 :safety(title,comments);分类 cs.CL、cs.AI
Comments ML Safety Workshop, NeurIPS 2022
专题命中 AI治理与伦理 :trustworthy(title,comments);分类 cs.AI、cs.CY
Comments 10 pages, 2 tables, pre-print approved for publication in the Special Issue Reflections on Responsible Research and Innovation for Trustworthy Autonomous Systems in the Journal of Responsible Technology
伦理决策头:基于人类反馈的强化学习在自动驾驶中实现规范伦理
机构 * Stony Brook University(石溪大学)
专题命中 AI治理与伦理 :RLHF(abstract,abstract_cn);safety(abstract);分类 cs.LG
AI总结 本文提出伦理决策头(EDH)框架,结合PPO与人类偏好奖励模型,在CARLA仿真中训练自动驾驶智能体,发现人类对自动驾驶伦理的理论规定与实践奖励存在差异。
立场:AI锁定正在发生,我们必须做好准备
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI
AI总结 该研究指出AI锁定是被低估的AI安全风险,分析其在个体、社会、国家层面的形成与升级机制,提出需提前应对以维护自主权与国家安全。
Comments ICML 2026 Position Track Spotlight
LLaMA 3.1-8B-Instruct中的框架条件化道德计算:伦理推理的机械可解释性审计
机构 * KD Consulting, CA, USA(KD咨询公司,美国加利福尼亚州) ; New York University, NY, USA(纽约大学,美国纽约州)
专题命中 AI治理与伦理 :RLHF(abstract,abstract_cn);alignment(abstract);分类 cs.AI
AI总结 通过机械可解释性平台分析LLaMA 3.1-8B-Instruct在54个道德提示上的内部计算,发现情境锚定效应:领域特定表示主导激活列表顶部,模型道德能力恒定但显著性高度依赖于提示选择的解释框架。
Comments 47 pages, 10 figures
AI可信性:可验证AI治理的新范式
机构 * AI Integrity Organization (AIO)(人工智能诚信组织(AIO))
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI
AI总结 本文提出AI可信性概念,旨在通过保护AI系统中的权威堆栈,确保推理过程可验证,不同于现有AI伦理、安全和对齐范式。
Comments 13 pages, 8 tables
ARYA:一种受物理约束的可组合且确定性世界模型架构
机构 * ARYA Labs(ARYA实验室)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI
AI总结 本文提出ARYA,一种基于五项原则的可组合、受物理约束且确定性世界模型架构,通过层级系统实现高效能与计算效率的平衡,展示其在六个基准测试中的卓越表现。
可靠且负责任的基础模型:全面综述
专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY
AI总结 本文综述了基础模型的可靠和负责任发展,探讨了偏见、安全、不确定性等关键问题,并提出了未来研究方向。
Comments TMLR camera-ready version
大语言模型部署中的伦理风险:医疗伦理“ Jailbreak”评估
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CY
AI总结 本文评估了大语言模型在医疗伦理领域面临的安全风险,发现七种主流模型在对抗测试中表现不一,其中Claude-Sonnet-4-Reasoning表现最稳健,而其他五种模型几乎全部失效。
通过人类之眼审视人工智能:在机器心理学中探究认知理论
机构 * Heritage Institute of Technology(赫里蒂奇理工学院)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI
AI总结 本文通过四个心理学框架研究LLMs的认知模式,发现其在叙述生成、框架偏差、道德判断和自我矛盾等方面表现出与人类相似但受训练数据影响的行为特征。
Comments Accepted to IJCNLP-AACL 2025 Student Research Workshop
从准确度到影响:用于对齐工程架构与影响理论的影响力驱动AI框架(IDAIF)
专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);trustworthy(abstract);分类 cs.AI
AI总结 IDAIF通过整合影响理论与AI架构,提供了一种以影响为中心的AI开发框架,旨在提升AI系统的伦理性和社会价值。
机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)科学与工程学院) ; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院) ; Shenzhen Institute of Artificial Intelligence and Robotics for Society, China(深圳人工智能与机器人研究院)
专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL
机构 * Yale University(耶鲁大学) ; National Library of Medicine, National Institutes of Health(国家医学图书馆,国立卫生研究院) ; Mila-Quebec AI Institute(魁北克AI研究所) ; Shanghai Jiao Tong University(上海交通大学) ; OPPO Research Institute(OPPO研究院) ; Reichman University(里奇曼大学)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY
机构 * Sebastian Dumbrava(独立研究者)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI
Comments 18 pages
专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.CL
专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY
Comments Published in Science: https://www.science.org/doi/10.1126/science.adn0117
专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY