arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2504.16148 2025-12-03 cs.CY cs.AI 62%

Towards responsible AI for education: Hybrid human-AI to confront the Elephant in the room

迈向负责任的教育AI:混合人类-人工智能以应对教育领域的关键问题

Danial Hooshyar, Gustav Šír, Yeongwook Yang, Eve Kikas, Raija Hämäläinen, Tommi Kärkkäinen, Dragan Gašević, Roger Azevedo

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

AI总结 本文探讨教育AI中的关键问题,提出神经符号AI作为解决这些问题的混合方法,以实现负责任的AI系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02265 2025-12-03 cs.LG cs.CY 62%

The Effect of Enforcing Fairness on Reshaping Explanations in Machine Learning Models

在机器学习模型中强制公平性对解释重塑的影响

Joshua Wolff Anderson, Shyam Visweswaran

机构 * Intelligent Systems Program, University of Pittsburgh(1 智能系统计划,匹兹堡大学) Department of Biomedical Informatics, University of Pittsburgh(2 生物医学信息学系,匹兹堡大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

AI总结 本研究探讨了在医疗机器学习中通过偏见缓解技术提高公平性如何影响基于Shapley的特征排名,发现公平性提升可能改变特征重要性排名,强调了在模型评估中需综合考虑准确性、公平性和可解释性。

Comments 10 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02058 2025-12-03 cs.CY cs.CL 62%

Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings

在低资源环境下,LLM生成的人设与人类感知的错位

Tabia Tanzin Prama, Christopher M. Danforth, Peter Sheridan Dodds

机构 * Computational Story Lab(计算故事实验室) Vermont Complex Systems Institute(佛罗里达复杂系统研究所) Vermont Advanced Computing Center(佛罗里达高级计算中心) Department of Mathematics and Statistics(数学与统计学系) Department of Computer Science University of Vermont(佛罗里达大学计算机科学系) Santa Fe Institute(圣菲研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 本研究发现,在低资源环境中,LLM生成的人设在共情和可信度方面显著劣于人类,需通过现实数据验证以确保其可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00742 2025-12-02 cs.CY cs.AI 62%

On the Regulatory Potential of User Interfaces for AI Agent Governance

关于用户界面在AI代理治理中的调节潜力

K. J. Kevin Feng, Tae Soo Kim, Rock Yuren Pang, Faria Huq, Tal August, Amy X. Zhang

机构 * University of Washington(华盛顿大学) KAIST(韩国科学技术院) Carnegie Mellon University(卡内基梅隆大学) UIUC(伊利诺伊大学香槟分校)

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI、cs.CY

AI总结 本文提出通过调节AI代理的用户界面来增强透明性和行为规范,从而在系统和基础设施层面实现治理。

Comments RegML workshop at NeurIPS 2025 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00461 2025-12-02 cs.CY cs.CL 62%

Whose Personae? Synthetic Persona Experiments in LLM Research and Pathways to Transparency

谁的人设?LLM研究中的人设实验及透明化路径

Jan Batzner, Volker Stocker, Bingjun Tang, Anusha Natarajan, Qinhao Chen, Stefan Schmid, Gjergji Kasneci

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 本文探讨了LLM研究中合成人设实验的代表性与生态效度问题,提出透明化检查表以提升评估的严谨性和实证性。

Comments Published at AAAI/ACM AIES 2025. Presented at NeurIPS 2025 Workshop Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

Journal ref Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8(1), 2025, 343-354

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20680 2025-11-27 cs.CL cs.AI 62%

Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

大语言模型的认知偏差影响临床肿瘤学笔记的解读

Matthew W. Kenaston, Umair Ayub, Mihir Parmar, Muhammad Umair Anjum, Syed Arsalan Ahmed Naqvi, Priya Kumar, Samarth Rawal, Aadel A. Chaudhuri, Yousef Zakharia, Elizabeth I. Heath, Tanios S. Bekaii-Saab, Cui Tao, Eliezer M. Van Allen, Ben Zhou, YooJung Choi, Chitta Baral, Irbaz Bin Riaz

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本研究揭示了大语言模型在肿瘤学笔记解读中因推理缺陷导致的临床安全隐患,并提出了一种可推广的推理错误分类框架。

Comments 24 pages, 6 figures, 1 supplementary figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18182 2025-11-25 cs.CY cs.AI 62%

The Workflow as Medium: A Framework for Navigating Human-AI Co-Creation

流程作为媒介:一种导航人机协同创作的框架

Lee Ackerman

机构 * Media University of Applied Sciences(应用科学媒体大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出创意智能循环框架,通过图文小说探讨人工智能在人机协同创作中的伦理与治理挑战,推动AI素养提升。

Comments 57 pages, 13 images, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04695 2025-11-21 cs.AI cs.CE cs.ET cs.LG 62%

Bridging the Gap in XAI-Why Reliable Metrics Matter for Explainability and Compliance

弥合XAI差距:可靠度量在可解释性与合规性中的重要性

Pratinav Seth, Vinay Kumar Sankarapu

机构 * Lexsi Labs(Lexsi实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出以度量治理的范式,通过标准化指标提升AI系统的可解释性和合规性,防止对齐造假,构建持续的AI保证流程。

Comments Accepted at first EurIPS Workshop on Private AI Governance

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14606 2025-11-19 cs.CL cs.LG 62%

Bridging Human and Model Perspectives: A Comparative Analysis of Political Bias Detection in News Media Using Large Language Models

Shreya Adrita Banik, Niaz Nafi Rahman, Tahsina Moiukh, Farig Sadeque

机构 * Department of Computer Science and Engineering, BRAC University(计算机科学与工程系,BRAC大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18708 2025-11-19 cs.MA cs.AI cs.LG 62%

Skill-Aligned Fairness in Multi-Agent Learning for Collaboration in Healthcare

Promise Osaine Ekpo, Brian La, Thomas Wiener, Saesha Agarwal, Arshia Agrawal, Gonzalo Gonzalez-Pumariega, Lekan P. Molu, Angelique Taylor

机构 * Cornell Tech(康奈尔科技) Microsoft Research(微软研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11790 2025-11-18 cs.CY cs.AI 62%

Differences in the Moral Foundations of Large Language Models

Peter Kirgis

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10089 2025-11-18 cs.LG cs.AI 62%

T2IBias: Uncovering Societal Bias Encoded in the Latent Space of Text-to-Image Generative Models

Abu Sufian, Cosimo Distante, Marco Leo, Hanan Salam

机构 * National Research Council of Italy - Institute of Applied Sciences

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments This manuscript has been accepted for presentation in the First Interdisciplinary Workshop on Responsible AI for Value Creation. Dec 1, Copenhagen. The final version will be submitted for inclusion in a Springer LNCS Volume. (The paper is 15 pages with 7 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05927 2025-11-13 cs.CY cs.AI econ.GN q-fin.EC 62%

Artificial intelligence and the Gulf Cooperation Council workforce adapting to the future of work

Mohammad Rashed Albous, Melodena Stephens, Odeh Rashed Al-Jayyousi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Journal ref Humanit Soc Sci Commun 12, 1649 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08567 2025-11-12 cs.LG cs.AI 62%

The Path Not Taken: RLVR Provably Learns Off the Principals

Hanqing Zhu, Zhenyu Zhang, Hanxian Huang, DiJia Su, Zechun Liu, Jiawei Zhao, Igor Fedorov, Hamed Pirsiavash, Zhizhou Sha, Jinwon Lee, David Z. Pan, Zhangyang Wang, Yuandong Tian, Kai Sheng Tai

机构 * Meta AI The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments Preliminary version accepted as a spotlight in NeurIPS 2025 Workshop on Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08082 2025-11-12 cs.AI cs.LG econ.GN q-fin.EC 62%

Prudential Reliability of Large Language Models in Reinsurance: Governance, Assurance, and Capital Efficiency

Stella C. Dong

机构 * Reinsurance Analytics(再保险分析)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments 48 pages, 9 figures, 5 tables. Submitted to the Journal of Risk and Insurance (JRI), November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07803 2025-11-12 cs.CY cs.AI 62%

Judging by the Rules: Compliance-Aligned Framework for Modern Slavery Statement Monitoring

Wenhao Xu, Akshatha Arodi, Jian-Yun Nie, Arsene Fansi Tchango

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments To appear at AAAI-26 (Social Impact Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04698 2025-11-11 cs.CL cs.AI 62%

multiMentalRoBERTa: A Fine-tuned Multiclass Classifier for Mental Health Disorder

K M Sajjadul Islam, John Fields, Praveen Madiraju

机构 * Marquette University, WI, USA(马quette大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted in IEEE Big Data, 8-11 December, 2025 @ Macau SAR, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03262 2025-11-11 cs.CL cs.LG 62%

REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Jian Hu, Jason Klein Liu, Haotian Xu, Wei Shen

专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.CL、cs.LG

Comments refactor

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05766 2025-11-11 cs.AI cs.CL econ.GN q-fin.EC 62%

Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs

Felipe Valencia-Clavijo

机构 * Dataplicada

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03980 2025-11-07 cs.AI cs.CL 62%

LLMs and Cultural Values: the Impact of Prompt Language and Explicit Cultural Framing

Bram Bulté, Ayla Rigouts Terryn

机构 * Brussels Centre for Language Studies, Vrije Universiteit Brussel(布鲁塞尔语言研究中心,布鲁塞尔自由大学) Université de Montréal & Mila - Quebec AI Institute(蒙特利尔大学及魁北克人工智能研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Preprint under review at Computational Linguistics. Accepted with minor revisions (10/10/2025); second round

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17228 2025-11-06 cs.CY cs.AI 62%

Survey on AI Ethics: A Socio-technical Perspective

Dave Mbiazi, Meghana Bhange, Maryam Babaei, Ivaxi Sheth, Patrik Kenfack, Samira Ebrahimi Kahou

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments Updated to the peer-reviewed version accepted and published in Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)

Journal ref Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10603 2025-10-31 cs.CY cs.AI 62%

Toward a Public and Secure Generative AI: A Comparative Analysis of Open and Closed LLMs

Jorge Machado

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25218 2025-10-30 cs.CY cs.AI 62%

Human Resilience in the AI Era -- What Machines Can't Replace

Shaoshan Liu, Anina Schwarzenbach, Yiyu Shi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22823 2025-10-28 cs.CL cs.AI 62%

Cross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLP

Poli Nemkova, Amrit Adhikari, Matthew Pearson, Vamsi Krishna Sadu, Mark V. Albert

机构 * University of North Texas(北卡罗来纳大学达顿分校) Davidson College(戴维森学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16836 2025-10-27 cs.CL cs.AI 62%

Misspellings in Natural Language Processing: A survey

Gianluca Sperduti, Alejandro Moreo

机构 * Istituto di Scienza e Tecnologie dell’Informazione, Consiglio Nazionale delle Ricerche(信息科学与技术研究所,国家研究理事会)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20782 2025-10-24 cs.CL cs.AI 62%

A Use-Case Specific Dataset for Measuring Dimensions of Responsible Performance in LLM-generated Text

Alicia Sagae, Chia-Jung Lee, Sandeep Avula, Brandon Dang, Vanessa Murdock

机构 * AWS Responsible AI(AWS负责任人工智能)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments 24 pages with 3 figures, to appear in Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19733 2025-10-24 cs.CL cs.LG 62%

Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning

M. H. I. Abdalla, Zhipin Wang, Christian Frey, Steffen Eger, Josif Grabocka

机构 * Department of Computer Science University of Technology Nuremberg(计算机科学系图腾技术大学纽伦堡)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18193 2025-10-23 cs.AI cs.CV cs.LG stat.ML 62%

FST.ai 2.0: An Explainable AI Ecosystem for Fair, Fast, and Inclusive Decision-Making in Olympic and Paralympic Taekwondo

Keivan Shariatmadar, Ahmad Osman, Ramin Ray, Kisam Kim

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 23 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14106 2025-10-17 cs.AI cs.CL cs.GT 62%

Generating Fair Consensus Statements with Social Choice on Token-Level MDPs

Carter Blair, Kate Larson

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08236 2025-10-17 cs.LG cs.AI 62%

The Hidden Bias: A Study on Explicit and Implicit Political Stereotypes in Large Language Models

Konrad Löhr, Shuzhou Yuan, Michael Färber

机构 * Technische Universität Dresden(德累斯顿技术大学) Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI)(可扩展数据与人工智能研究中心(ScaDS.AI))

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏