arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1852 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1852 篇

2508.04071 2025-08-07 cs.LG 57%

Adversarial Fair Multi-View Clustering

Mudi Jiang, Jiahui Zhou, Lianyu Hu, Xinying Liu, Zengyou He, Zhikui Chen

机构 * School of Software, Dalian University of Technology(大连理工大学软件学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17945 2025-08-07 cs.CL 57%

Assessing Agentic Large Language Models in Multilingual National Bias

Qianying Liu, Katrina Qiyao Wang, Fei Cheng, Sadao Kurohashi

机构 * National Institute of Informatics, Japan(日本信息机构国家研究所) University of Wisconsin—Madison, USA(美国威斯康星大学麦迪逊分校) Kyoto University, Japan(日本京都大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted to ACL 2025 Findings. 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10945 2025-08-07 cs.LG stat.ML 57%

Gradient-Based Multi-Objective Deep Learning: Algorithms, Theories, Applications, and Beyond

Weiyu Chen, Baijiong Lin, Xiaoyuan Zhang, Xi Lin, Han Zhao, Qingfu Zhang, James T. Kwok

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) City University of Hong Kong(香港城市大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00847 2025-08-05 cs.HC cs.CY 57%

GPT Chatbots for Alleviating Anxiety and Depression: A Pilot Randomized Controlled Trial with Afghan Women

Sofia Sahab, Jawad Haqbeen, Diksha Sapkota, Takayuki Ito

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23454 2025-08-04 cs.HC cs.CY cs.ET cs.GR q-bio.NC 57%

Breaking the mould of Social Mixed Reality - State-of-the-Art and Glossary

Marta Bieńkiewicz, Julia Ayache, Panayiotis Charalambous, Cristina Becchio, Marco Corragio, Bertram Taetz, Francesco De Lellis, Antonio Grotta, Anna Server, Daniel Rammer, Richard Kulpa, Franck Multon, Azucena Garcia-Palacios, Jessica Sutherland, Kathleen Bryson, Stéphane Donikian, Didier Stricker, Benoît Bardy

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21929 2025-07-30 cs.AI 57%

Libra: Large Chinese-based Safeguard for AI Content

Ziyang Chen, Huimu Yu, Xing Wu, Dongqin Liu, Songlin Hu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21319 2025-07-30 cs.CL 57%

Do Large Language Models Understand Morality Across Cultures?

Hadi Mohammadi, Yasmeen F. S. S. Meijer, Efthymia Papadopoulou, Ayoub Bagheri

机构 * Department of Methodology and Statistics, Utrecht University, The Netherlands(方法论与统计学系,乌特雷赫特大学,荷兰)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19962 2025-07-29 cs.CL 57%

KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models

Seorin Kim, Dongyoung Lee, Jaejin Lee

机构 * Dept. of Data Science, Seoul National University(数据科学系,首尔国立大学) Dept. of Computer Science and Engineering, Seoul National University(计算机科学与工程系,首尔国立大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13138 2025-07-29 cs.CL 57%

Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation

Hadi Mohammadi, Tina Shahedi, Pablo Mosteiro, Massimo Poesio, Ayoub Bagheri, Anastasia Giachanou

机构 * Department of Methodology and Statistics, Utrecht University, The Netherlands(方法论与统计学系,乌特列支大学,荷兰) Department of Information and Computing Sciences, Utrecht University, The Netherlands(信息与计算科学系,乌特列支大学,荷兰) Queen Mary University of London, London, United Kingdom(伦敦女王玛丽大学,伦敦,英国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13175 2025-07-29 cs.AI 57%

Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era

Matthew E. Brophy

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 42 pages. Supplementary material included at end of article

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17368 2025-07-24 cs.LG 57%

ViRN: Variational Inference and Distribution Trilateration for Long-Tailed Continual Representation Learning

Hao Dai, Chong Tang, Jagmohan Chauhan

机构 * Department of Computer Science, UCL Centre for Artificial Intelligence, University College London, London, UK(计算机科学系,UCL人工智能中心,伦敦大学学院,伦敦,英国) University of Southampton, Southampton, UK(南安普顿大学,南安普顿,英国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13294 2025-07-24 cs.CY cs.HC 57%

The "Who", "What", and "How" of Responsible AI Governance: A Systematic Review and Meta-Analysis of (Actor, Stage)-Specific Tools

Blaine Kuehnert, Rachel M. Kim, Jodi Forlizzi, Hoda Heidari

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments Accepted to ACM Conference on Fairness, Accountability, and Transparency 2025. 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14226 2025-07-23 cs.CY 57%

Mapping the Parasocial AI Market: User Trends, Engagement and Risks

Zilan Qian, Mari Izumikawa, Fiona Lodge, Angelo Leone

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 17 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14332 2025-07-22 cs.LG 57%

Development and Deployment of Hybrid ML Models for Critical Heat Flux Prediction in Annulus Geometries

Aidan Furlong, Xingang Zhao, Robert Salko, Xu Wu

机构 * Department of Nuclear Engineering, North Carolina State University(核工程系,北卡罗来纳州立大学) Department of Nuclear Engineering, University of Tennessee(核工程系,田纳西大学) Nuclear Energy and Fuel Cycle Division, Oak Ridge National Laboratory(核能与燃料循环 division,橡树岭国家实验室)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments Accepted for inclusion in Transactions of the American Nuclear Society for the 2025 ANS Winter Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13616 2025-07-21 cs.HC cs.CY cs.ET cs.IT cs.MA math.IT 57%

From Firms to Computation: AI Governance and the Evolution of Institutions

Michael S. Harre

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments 44 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19232 2025-07-18 cs.IR cs.AI 57%

LLM-RecG: A Semantic Bias-Aware Framework for Zero-Shot Sequential Recommendation

Yunzhe Li, Junting Wang, Hari Sundaram, Zhining Liu

机构 * University of Illinois, Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 10 pages, Recsys'25 Spotlight Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11344 2025-07-16 cs.LG 57%

Guiding LLM Decision-Making with Fairness Reward Models

Zara Hall, Melanie Subbiah, Thomas P Zollo, Kathleen McKeown, Richard Zemel

机构 * Columbia University(哥伦比亚大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10861 2025-07-16 cs.LG 57%

Visually grounded emotion regulation via diffusion models and user-driven reappraisal

Edoardo Pinzuti, Oliver Tüscher, André Ferreira Castro

机构 * Leibniz Institute for Resilience Research(莱比锡韧性研究所) University Medicine Halle (Saale) of the Martin Luther University Halle-Wittenberg (MLU)(马尔堡-哈雷大学哈雷-萨勒医学院) German Center for Mental Health (DZPG)(德国心理健康中心) School of Life Sciences, Technical University of Munich(慕尼黑技术大学生命科学学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08979 2025-07-15 cs.CV cs.LG 57%

PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding Projection

Mahdiyar Molahasani, Azadeh Motamedi, Michael Greenspan, Il-Min Kim, Ali Etemad

机构 * Queen’s University(皇后大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08594 2025-07-14 cs.SE cs.AI cs.HC 57%

Generating Proto-Personas through Prompt Engineering: A Case Study on Efficiency, Effectiveness and Empathy

Fernando Ayach, Vitor Lameirão, Raul Leão, Jerfferson Felizardo, Rafael Sobrinho, Vanessa Borges, Patrícia Matsubara, Awdren Fontão

机构 * Faculty of Computing - Federal University of Mato Grosso do Sul(计算机学院 - 莫扎尔河大省联邦大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 12 pages; 2 figures; Preprint with the original submission accepted for publication at 39th Brazilian Symposium on Software Engineering (SBES)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07604 2025-07-11 cs.LG q-bio.QM q-bio.TO 57%

Synthetic MC via Biological Transmitters: Therapeutic Modulation of the Gut-Brain Axis

Sebastian Lotter, Elisabeth Mohr, Andrina Rutsch, Lukas Brand, Francesca Ronchi, Laura Díaz-Marugán

机构 * Charité – Universitätsmedizin Berlin, Humboldt-Universität zu Berlin, Berlin Institute of Health (BIH), Berlin, Germany(柏林查理医院、洪堡-柏林大学、柏林健康研究所(BIH)、柏林)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05866 2025-07-09 cs.CY 57%

Understanding support for AI regulation: A Bayesian network perspective

Andrea Cremaschi, Dae-Jin Lee, Manuele Leonelli

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03458 2025-07-08 cs.CV cs.AI 57%

Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach

Leyan Xue, Zongbo Han, Guangyu Wang, Qinghua Hu, Mingyue Cheng, Changqing Zhang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22372 2025-06-30 cs.IR cs.CL 57%

Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement

Maryam Mousavian, Zahra Abbasiantaeb, Mohammad Aliannejadi, Fabio Crestani

机构 * Università della Svizzera italiana \& University of Amsterdam University of Amsterdam The Netherland Unviersity of Amsterdam The Netherland University of Amsterdam

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted by ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20743 2025-06-27 cs.LG cs.CE 57%

A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools

Minh-Hao Van, Prateek Verma, Chen Zhao, Xintao Wu

机构 * Department of EECS University of Arkansas Fayetteville(电子工程与计算机科学系美国阿肯色大学弗莱维尔分校) Department of CS Baylor University Waco(计算机科学系贝勒大学沃斯堡)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18732 2025-06-24 cs.LG 57%

Towards Group Fairness with Multiple Sensitive Attributes in Federated Foundation Models

Yuning Yang, Han Yu, Tianrun Gao, Xiaodong Xu, Guangyu Wang

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18802 2025-06-24 cs.CL 57%

Language Models Grow Less Humanlike beyond Phase Transition

Tatsuya Aoyama, Ethan Wilcox

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted to ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04345 2025-06-23 cs.CY 57%

Build Agent Advocates, Not Platform Agents

Sayash Kapoor, Noam Kolt, Seth Lazar

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Accepted to ICML 2025 position paper track

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04686 2025-06-19 cs.AI 57%

Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization

Zelai Xu, Wanjun Gu, Chao Yu, Yi Wu, Yu Wang

机构 * Tsinghua University, Beijing, China(清华大学) Beijing Zhongguancun Academy, Beijing, China(北京中关村学院) Shanghai Qi Zhi Institute, Shanghai, China(上海启智研究所)

专题命中 AI治理与伦理 :DPO(abstract);分类 cs.AI

Comments Published in ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12527 2025-06-17 cs.CL 57%

Detection, Classification, and Mitigation of Gender Bias in Large Language Models

Xiaoqing Cheng, Hongying Zan, Lulu Kong, Jinwang Song, Min Peng

机构 * Zhengzhou University(郑州大学) Wuhan University(武汉大学)

专题命中 AI治理与伦理 :DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏