Meta-learners' learning dynamics are unlike learners'
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
Comments 26 pages, 23 figures
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
Comments 26 pages, 23 figures
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
Comments 5,240 word texts, 3 tables, 14 figures. Transportation Research Record: Journal of the Transportation Research Board, 2019
专题命中 其他安全 :safety(abstract);分类 cs.CY、cs.LG
Journal ref npj Digital Medicine 1:36 (2018)
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
Comments 6 pages; 1 figure; title, abstract updated; new experimental results
Journal ref Proceedings of the AAAI 2019 Workshop on SafeAI
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
Comments NIPS 2018. Version 2 contains more experimental data including best hyperparameters found
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG
Comments 9 pages
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY
Comments This paper will appear in Communications of the ACM (December 2018, vol 61, no. 12)
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
Comments Published at DECOR / ICDE 2018. Extended version accepted at SIGIR 2018, available here: arXiv:1804.11146
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
Comments 22 pages. To be submitted to IEEE Transactions on Software Engineering
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
Comments 6 pages, accepted by NAACL2016 short paper
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG
Comments The paper is accepted at FSE 2016 (the 24th ACM SIGSOFT International Symposium on the Foundations of Software Engineering)
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
Journal ref Journal Of Artificial Intelligence Research, Volume 32, pages 793-824, 2008
通过潜在视角评估LLM中的多元主义
机构 * University of Helsinki(赫尔辛基大学) ; ETH Zurich(苏黎世联邦理工学院)
专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL
AI总结 提出一种领域无关的多层无监督框架,从LLM生成文本中提取潜在视角,评估多元主义差距,发现稀有视角仍被不成比例地低估。
Comments Pluralistic Alignment Workshop @ ICML 2026
IGBT模块焊层退化与温度监测的虚拟传感
机构 * Silicon Austria Labs GmbH(硅 Austria 实验室)
专题命中 其他安全 :safety(abstract,journal_ref);分类 cs.LG
AI总结 本文利用机器学习虚拟传感技术,通过有限物理传感器估计IGBT模块焊层退化状态及温度分布,实现高精度退化区域估计和表面温度复现。
Comments Andrea Urgolo and Monika Stipsitz contributed equally to this work
Journal ref 2025 9th International Conference on System Reliability and Safety (ICSRS), Turin, Italy, 2025, pp. 538-547
人工智能现象学:理解跨时代的以人为本的AI体验
机构 * ETH Zürich(苏黎世联邦理工学院) ; University of Bergen(卑尔根大学) ; Stanford University(斯坦福大学)
专题命中 其他安全 :alignment(abstract,comments);分类 cs.AI
AI总结 本文提出人工智能现象学,通过研究用户与AI交互的主观体验,促进双向人机对齐,并提供可重复的方法论工具。
Comments This is an accepted workshop paper at CHI '26, "W37: Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures", or https://bialign-workshop.github.io/2026/cfp
为实践者设计生成式AI:探索与创意实践对齐的交互方法
机构 * LISN Université Paris-Saclay, CNRS, Inria(LISN 巴黎-萨克雷大学,法国国家科学研究中心,法国国家信息与自动化技术研究所) ; UMR 9189 CRIStAL Univ. Lille, Inria, CNRS, Centrale Lille(UMR 9189 CRIStAL 利摩日大学,法国国家信息与自动化技术研究所,法国国家科学研究中心,利摩日中央理工大学)
专题命中 其他安全 :alignment(abstract,comments);分类 cs.AI
AI总结 本文提出三种交互方法,帮助设计师在不同阶段引导生成式AI与创意实践对齐,强调动态协商与主动/被动角色的适应性。
Comments Accepted to ACM CHI 2026 Workshop on Bidirectional Human-AI Alignment
机构 * Novus Technologies ; MIT(麻省理工学院) ; Harvard University(哈佛大学)
专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL
Comments ICLR 2024 Workshop on Representational Alignment (Re-Align) Camera Ready
机构 * Digitale Schiene Deutschland, DB InfraGO AG(Digitale Schiene Deutschland,DB InfraGO AG) ; Communication Systems Group, Technische Universität Berlin(通信系统组,技术大学柏林)
专题命中 其他安全 :safety(abstract,comments);分类 cs.LG
Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in Computer Safety, Reliability and Security - SAFECOMP 2024 Workshops - DECSoS, SASSUR, TOASTS, and WAISE, and is available online at https://doi.org/10.1007/978-3-031-68738-9_29
专题命中 其他安全 :alignment(abstract,comments);分类 cs.AI
Comments This paper will appear at ICLR 2025 Workshop on Bidirectional Human-AI Alignment
专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL
Comments 9 pages, 8 figures, Accepted to AAAI 2025 Main Conference (AI Alignment Track)
专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL
Comments To appear at the ICLR 2024 Workshop on Representational Alignment (Re-Align)
专题命中 其他安全 :safety(abstract,comments);分类 cs.AI
Comments 18 pages, 12 figures, to appear in the proceedings of the 2023 Safety-critical Systems Symposium (SSS '23), York, UK
专题命中 其他安全 :safety(abstract,comments);分类 cs.LG
Comments Accepted in The 6th International Conference on System Reliability and Safety (ICSRS) 2022
专题命中 其他安全 :safety(abstract,comments);分类 cs.AI
Comments submitted version; appeared at: International Conference on Computer Safety, Reliability, and Security. Springer, Cham, 2021
专题命中 其他安全 :safety(abstract,journal_ref);分类 cs.AI
Comments http://ceur-ws.org/Vol-2560/paper44.pdf
Journal ref Proceedings of the Workshop on Artificial Intelligence Safety (SafeAI 2020) co-located with 34th AAAI Conference on Artificial Intelligence (AAAI 2020), New York, USA, Feb 7, 2020