arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2601.19280 2026-01-28 cs.LG cs.AI cs.CL 67%

Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning

基于组分布鲁棒优化的强化学习用于大语言模型推理

Kishan Panaganti, Zhenwen Liang, Wenhao Yu, Haitao Mi, Dong Yu

机构 * Tencent AI Lab in Bellevue WA USA(腾讯AI实验室(西雅图华盛顿州))

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出多对手组分布鲁棒优化框架,通过动态调整训练分布提升大语言模型推理性能,实现训练后精度提升10.6%和10.1%。

Comments Keywords: Large Language Models, Reasoning Models, Reinforcement Learning, Distributionally Robust Optimization, GRPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17217 2026-01-15 cs.CL cs.AI cs.CY 67%

Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs

通过促进大语言模型的探索性思维来缓解性别偏见

Kangda Wei, Hasnat Md Abdullah, Ruihong Huang

专题命中 AI治理与伦理 :DPO(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 通过生成性别中性故事对并利用直接偏好优化,该研究旨在减少大语言模型中的性别偏见,同时保持模型能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03203 2026-01-07 cs.LG cs.AI cs.CY 67%

Counterfactual Fairness with Graph Uncertainty

基于图不确定性的反事实公平性

Davi Valério, Chrysoula Zerva, Mariana Pinto, Ricardo Santos, André Carreiro

机构 * Instituto Superior Técnico(里斯本技术高等学院) Instituto de Telecomunicações(电信研究所) Fraunhofer Portugal AICOS(弗劳恩霍夫葡萄牙AICOS研究所)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文提出CF-GU方法,通过整合因果图的不确定性,提升反事实公平性评估的鲁棒性和准确性。

Comments Peer reviewed pre-print. Presented at the BIAS 2025 Workshop at ECML PKDD

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11671 2025-12-24 cs.AI cs.CY cs.LG econ.GN q-fin.EC 67%

Computational Basis of LLM's Decision Making in Social Simulation

大语言模型在社会模拟中的决策机制计算基础

Ji Ma

机构 * LBJ School of Public Affairs, University of Texas at Austin(德克萨斯大学奥斯汀分校公共事务学院LBJ学院) Gradel Institute of Charity, New College, University of Oxford(牛津大学格拉德尔慈善研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本研究通过独裁者游戏探索LLM内部表示的变量变化,揭示社会概念在Transformer模型中的编码机制,为社会模拟和AI对齐提供新方法。

Comments Forthcoming: Sociological Methodology; USPTO patent pending

Journal ref Sociological Methodology, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15693 2025-12-16 cs.LG cs.AI cs.CL 67%

Beyond Benchmarks: On The False Promise of AI Regulation

超越基准:关于人工智能监管的虚假承诺

Gabriel Stanovsky, Renana Keydar, Gadi Perl, Eliya Habba

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Faculty of Law and(法学院) Center of Digital Humanities(数字人文中心)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出不依赖基准的人工智能监管框架,强调人工智能可解释性挑战对现有监管体系的制约,并呼吁跨学科合作解决这一关键问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01812 2025-11-27 cs.CY cs.AI cs.CL 67%

From Text to Multimodality: Exploring the Evolution and Impact of Large Language Models in Medical Practice

从文本到多模态:探索大型语言模型在医疗实践中的演变与影响

Qian Niu, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Lawrence KQ Yan, Yichao Zhang, Caitlyn Heqi Yin, Cheng Fei, Junyu Liu, Tianyang Wang, Yunze Wang, Silin Chen, Ming Liu, Benji Peng, Xinyuan Song, Ziyuan Qin, Riyang Bao, Zekun Jiang

机构 * Kyoto University(京都大学) Georgia Institute of Technology(佐治亚理工学院) National Taiwan Normal University(台湾师范大学) Indiana University(印第安纳大学) Hong Kong University of Science(香港科学大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Cornell University(康奈尔大学) University of Liverpool(利物浦大学) University of Edinburgh(爱丁堡大学) Zhejiang University(浙江大学) Purdue University(Purdue 大学) Emory University, Atlanta, GA, USA(埃默里大学) West China Biomedical Big Data Center, West China Hospital, Sichuan University, Chengdu, China(西京生物大数据中心,四川大学西京医院,成都,中国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文探讨了多模态大型语言模型在医疗实践中的发展与影响,分析其在医疗影像、临床决策支持等领域的应用及面临的挑战。

Comments 12 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19979 2025-11-26 cs.IR 67%

The 2nd Workshop on Human-Centered Recommender Systems

人类中心推荐系统研讨会第二届

Kaike Zhang, Jiakai Tang, Du Su, Shuchang Liu, Julian McAuley, Lina Yao, Qi Cao, Yue Feng, Fei Sun

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract)

AI总结 该研讨会旨在推动推荐系统从优化参与度向设计真正理解、参与和惠及人类的系统转变,探讨如何整合人类价值观以提升推荐系统的社会责任感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20432 2025-11-04 cs.AI cs.CY cs.GT cs.LG 67%

LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory

Jingru Jia, Zehua Yuan, Junhao Pan, Paul E. McNamara, Deming Chen

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15831 2025-10-23 cs.CL cs.AI cs.CY 67%

Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs

Vishnu Hari, Kalpana Panda, Srikant Panda, Amit Agarwal, Hitesh Laxmichand Patel

机构 * Birla Institute of Technology and Science (BITS)(巴拉·技术与科学学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09051 2025-10-13 cs.CL cs.AI cs.LG 67%

Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation

Muhammad Ali Shafique, Kanwal Mehreen, Muhammad Arham, Maaz Amjad, Sabur Butt, Hamza Farooq

机构 * University of British Columbia(不列颠哥伦比亚大学) Texas Tech University(德克萨斯技术大学) Institute for the Future of Education, Tecnológico de Monterrey(教育未来研究所,墨西哥蒙特雷技术学院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to the EMNLP 2025 Workshop on Multilingual Representation Learning (MRL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06105 2025-10-08 cs.AI cs.CY cs.HC cs.LG 67%

Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences

Batu El, James Zou

机构 * Stanford University(斯坦福大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10127 2025-10-07 cs.CL cs.AI cs.LG 67%

Population-Aligned Persona Generation for LLM-based Social Simulation

Zhengyu Hu, Jianxun Lian, Zheyuan Xiao, Max Xiong, Yuxuan Lei, Tianfu Wang, Kaize Ding, Ziang Xiao, Nicholas Jing Yuan, Xing Xie

机构 * HKUST(香港科技大学) Microsoft Research Asia(微软亚洲研究院) Duke University(杜克大学) Northwestern University(西北大学) Johns Hopkins University(约翰霍普金斯大学) Microsoft(微软)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01164 2025-10-02 cs.CL cs.AI cs.CY cs.HC 67%

Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare

Zhengliang Shi, Ruotian Ma, Jen-tse Huang, Xinbei Ma, Xingyu Chen, Mengru Wang, Qu Yang, Yue Wang, Fanghua Ye, Ziyang Chen, Shanyi Wang, Cixing Li, Wenxuan Wang, Zhaopeng Tu, Xiaolong Li, Zhaochun Ren, Linus

机构 * Tencent(腾讯)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03699 2025-09-23 cs.AI cs.CY cs.LG cs.MA q-bio.QM 67%

Enhancing Clinical Decision-Making: Integrating Multi-Agent Systems with Ethical AI Governance

Ying-Jung Chen, Ahmad Albarqawi, Chi-Sheng Chen

机构 * College of Computing(计算学院) Georgia Institute of Technology(佐治亚理工学院) University of Illinois(伊利诺伊大学) Neuro Industry Research(神经产业研究) Neuro Industry, Inc.(神经产业公司)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11648 2025-09-16 cs.CL cs.AI cs.CY 67%

EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI

Sai Kartheek Reddy Kasu

机构 * IIIT Dharwad(德瓦德理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17044 2025-09-16 cs.CY cs.AI cs.LG 67%

Approaches to Responsible Governance of GenAI in Organizations

Dhari Gandhi, Himanshu Joshi, Lucas Hartman, Shabnam Hassani

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) Western University(西部大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03034 2025-09-04 cs.LG cs.AI cs.CR cs.CV cs.CY 67%

Rethinking Data Protection in the (Generative) Artificial Intelligence Era

Yiming Li, Shuo Shao, Yu He, Junfeng Guo, Tianwei Zhang, Zhan Qin, Pin-Yu Chen, Michael Backes, Philip Torr, Dacheng Tao, Kui Ren

机构 * The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) Nanyang Technological University(南洋理工大学) University of Maryland(马里兰大学) IBM Research(IBM研究院) CISPA Helmholtz Center for Information Security(CISPA 欧洲信息安全部分) University of Oxford(牛津大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Perspective paper for a broader scientific audience. The first two authors contributed equally to this paper. 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19269 2025-08-28 cs.CY cs.AI cs.CL 67%

Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models

Ke Zhou, Marios Constantinides, Daniele Quercia

机构 * Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments This paper has been accepted in AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15860 2025-08-22 cs.CL cs.AI cs.LG 67%

Synthetic vs. Gold: The Role of LLM Generated Labels and Data in Cyberbullying Detection

Arefeh Kazemi, Sri Balaaji Natarajan Kalaivendan, Joachim Wagner, Hamza Qadeer, Kanishk Verma, Brian Davis

机构 * School of Computing, ADAPT Centre, Dublin City University, Dublin, Ireland(计算学院、ADAPT中心、都柏林城市大学、都柏林、爱尔兰)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07284 2025-08-12 cs.CL cs.AI cs.CY 67%

"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas

Junchen Ding, Penghao Jiang, Zihao Xu, Ziqi Ding, Yichen Zhu, Jiaojiao Jiang, Yuekang Li

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16663 2025-08-12 cs.LG cs.AI cs.CY cs.LO cs.SE 67%

Runtime Monitoring and Enforcement of Conditional Fairness in Generative AIs

Chih-Hong Cheng, Changshun Wu, Xingyu Zhao, Saddek Bensalem, Harald Ruess

机构 * Chalmers University of Technology, Sweden Carl von Ossietzky Universität Oldenburg, Germany Universit\'e Grenoble Alpes, France University of Warwick, United Kingdom CSX-AI, France SRI International, United States

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14339 2025-07-22 cs.CY cs.AI cs.HC cs.LG eess.SP 67%

Fiduciary AI for the Future of Brain-Technology Interactions

Abhishek Bhattacharjee, Jack Pilkington, Nita Farahany

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05442 2025-07-16 cs.AI cs.CY cs.HC cs.LG 67%

The Odyssey of the Fittest: Can Agents Survive and Still Be Good?

Dylan Waldner, Risto Miikkulainen

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted to CogSci 2025. Code can be found at https://github.com/dylanwaldner/BeGoodOrSurvive

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06434 2025-07-10 cs.CY cs.AI cs.LG 67%

Deprecating Benchmarks: Criteria and Framework

Ayrton San Joaquin, Rokas Gipiškis, Leon Staufer, Ariel Gil

机构 * AI Standards Lab(AI标准实验室) Technical University of Munich(慕尼黑技术大学) Institute of Data Science(数据科学研究所) Digital Technologies(数字技术) Trajectory Labs(轨迹实验室)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 10 pages, 1 table. Accepted to the ICML 2025 Technical AI Governance Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01630 2025-06-19 cs.LG cs.AI cs.CY 67%

Machine Learners Should Acknowledge the Legal Implications of Large Language Models as Personal Data

Henrik Nolte, Michèle Finck, Kristof Meding

机构 * University Tübingen(图宾根大学) CZS Institute for Artificial Intelligence and Law(法律与人工智能研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14230 2025-06-12 cs.CL cs.AI cs.CY 67%

Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing

Han Jiang, Xiaoyuan Yi, Zhihua Wei, Ziang Xiao, Shu Wang, Xing Xie

机构 * Tongji University(同济大学) Johns Hopkins University(约翰霍普金斯大学) University of California, Los Angeles(加州大学洛杉矶分校) Microsoft Research Asia(微软亚洲研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07045 2025-06-05 cs.CR cs.AI cs.CL cs.CY 67%

Scalable and Ethical Insider Threat Detection through Data Synthesis and Analysis by LLMs

Haywood Gelman, John D. Hastings

机构 * The Beacom College of Computer and Cyber Sciences(贝科姆计算机与网络科学学院) Dakota State University(达科他州立大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 6 pages, 0 figures, 8 tables

Journal ref 2025 IEEE 13th International Symposium on Digital Forensics and Security (ISDFS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21634 2025-05-01 cs.CY cs.AI cs.LG 67%

Quantitative Auditing of AI Fairness with Differentially Private Synthetic Data

Chih-Cheng Rex Yuan, Bow-Yaw Wang

机构 * Institute of Information Science, Academia Sinica(学术院信息研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05210 2025-04-08 cs.CY cs.AI cs.HC cs.LG 67%

A moving target in AI-assisted decision-making: Dataset shift, model updating, and the problem of update opacity

Joshua Hatherley

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Journal ref Ethics and Information Technology 27(2): 20 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15514 2025-04-08 cs.HC cs.AI cs.CL cs.CY cs.ET 67%

Superhuman Game AI Disclosure: Expertise and Context Moderate Effects on Trust and Fairness

Jaymari Chua, Chen Wang, Lina Yao

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏