arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-24 至 2026-02-24 共收录 9 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 9 篇

2602.18582 2026-02-24 cs.AI cs.CL cs.HC cs.LG 82%

Hierarchical Reward Design from Language: Enhancing Alignment of Agent Behavior with Human Specifications

基于语言的层次奖励设计:通过人类规范增强智能体行为对齐

Zhiqin Qian, Ryan Diaz, Sangwon Seo, Vaibhav Unhelkar

机构 * Rice University(里士大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出基于语言的层次奖励设计(HRDL)和语言到层次奖励(L2HR)方法,用于通过人类规范增强智能体行为对齐,提升AI任务完成效果和规范遵守程度。

Comments Extended version of an identically-titled paper accepted at AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19159 2026-02-24 cs.AI cs.CL cs.LG 75%

Beyond Behavioural Trade-Offs: Mechanistic Tracing of Pain-Pleasure Decisions in an LLM

超越行为权衡:在LLM中疼痛-愉悦决策的机制追溯

Francesca Bianco, Derek Shiller

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究揭示了LLM在疼痛-愉悦决策中的内部机制,通过机制追溯揭示了价值信号的表示和因果作用,为AI意识和福利的讨论提供了证据基础。

Comments 24 pages, 8+1 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13033 2026-02-24 cs.CY cs.AI cs.CE cs.CL cs.SI 67%

Buy versus Build an LLM: A Decision Framework for Governments

买还是建一个大语言模型:政府的决策框架

Jiahao Lu, Ziwei Xu, William Tjhi, Junnan Li, Antoine Bosselut, Pang Wei Koh, Mohan Kankanhalli

机构 * National University of Singapore(新加坡国立大学) AI Singapore(AI新加坡) Salesforce AI Research(Salesforce AI研究) EPFL(苏黎世联邦理工学院) University of Washington(华盛顿大学) Allen Institute for AI(人工智能研究院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出政府在大语言模型决策中应考虑主权、安全、成本等因素的框架,帮助确定购买或建设更适合其需求的方法。

Comments The short version of this document is published as an ACM TechBrief at https://dl.acm.org/doi/epdf/10.1145/3797946, and this document is published as an ACM Technology Policy Council white paper at https://www.acm.org/binaries/content/assets/public-policy/buildvsbuyai.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09905 2026-02-24 cs.LG cs.AI cs.CV 62%

PRISM: Diversifying Dataset Distillation by Decoupling Architectural Priors

PRISM: 通过解耦架构先验来多样化数据集蒸馏

Brian B. Moser, Shalini Sarode, Federico Raue, Stanislav Frolov, Krzysztof Adamkiewicz, Arundhati Shanbhag, Joachim Folz, Tobias C. Nauen, Andreas Dengel

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) RPTU Kaiserslautern-Landau(科隆大学(RPTU))

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 PRISM通过解耦架构先验提升数据集蒸馏的多样性与性能

Journal ref Transactions on Machine Learning Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18459 2026-02-24 cs.CY cs.AI 62%

From Bias Mitigation to Bias Negotiation: Governing Identity and Sociocultural Reasoning in Generative AI

从偏见缓解到偏见协商:在生成式AI中治理身份与社会文化推理

Zackary Okun Dunivin, Bingyi Han, John Bollenbocher

机构 * Institute for Social Science, University of Stuttgart(斯图加特大学社会科学研究所) Department of Computer Science, Saarland University(萨尔兰大学计算机科学系) Department of Linguistics, University of Texas Austin TX USA(德克萨斯大学语言学系) Department of Linguistics, University of Texas(德克萨斯大学语言学系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出偏见协商作为生成式AI治理身份与社会文化推理的新框架,通过实证研究探讨其在缓解偏见和提升社会公正中的作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19682 2026-02-24 cs.CY 57%

Beyond the Binary: A nuanced path for open-weight advanced AI

超越二元:开放权重高级AI的细致路径

Bengüsu Özcan, Alex Petropoulos, Max Reddel

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

AI总结 本文提出了一种基于安全评估的分层模型发布方法,旨在解决开放权重高级AI模型在安全性和监管方面的挑战。

Comments This publication was originally designed and optimised for web and published on cfg.eu. Minor formatting differences may appear in this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07754 2026-02-24 cs.AI cs.HC 57%

Humanizing AI Grading: Student-Centered Insights on Fairness, Trust, Consistency and Transparency

让AI评分更人性化:以学生为中心的公平性、信任、一致性与透明性洞察

Bahare Riahi, Viktoriia Storozhevykh, Veronica Catete

机构 * North Carolina State University(北卡罗来纳州立大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

AI总结 本研究通过比较AI与人工评分反馈,探讨学生对AI评分系统在公平性、信任、一致性与透明性方面的看法,并提出人本化AI的设计原则。

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18535 2026-02-24 cs.SD cs.AI 57%

Fairness-Aware Partial-label Domain Adaptation for Voice Classification of Parkinson's and ALS

面向语音分类的公平性意识部分标签领域适应

Arianna Francesconi, Zhixiang Dai, Arthur Stefano Moscheni, Himesh Morgan Perera Kanattage, Donato Cappetta, Fabio Rebecchi, Paolo Soda, Valerio Guarrasi, Rosa Sicilia, Mary-Anne Hartley

机构 * organization= School of Computer Communication Sciences, EPFL (\'Ecole polytechnique f\'ed\'erale de Lausanne) , city= Lausanne , country= Switzerland organization= Eustema S.p.A., Research Development Centre , city= Naples , country= Italy organization= UniCamillus-Saint Camillus International University of Health Sciences , city= Rome , country= Italy organization= Department of Diagnostics Intervention, Radiation Physics, Biomedical Engineering, Umeå University , city= Umeå , country= Sweden

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出了一种融合域泛化和对抗对齐的框架,用于在部分标签不匹配和公平性约束下实现帕金森病和肌萎缩侧索硬化症的统一语音分类。

Comments 7 pages, 1 figure. Submitted to Pattern Recognition Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18461 2026-02-24 cs.CY 57%

Toward Self-Driving Universities: Can Universities Drive Themselves with Agentic AI?

迈向自动驾驶大学:大学能否通过代理AI实现自我驱动?

Anis Koubaa

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本研究提出通过代理AI实现高等教育机构的自主性框架,旨在自动化行政、学术和质量保证流程,减少教师文书工作时间,提升教育质量和研究生产力。

Journal ref Springer Book: Next Generation AI-Driven Education - 2026

详情

展开后加载摘要…

URL PDF HTML 收藏