arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1852 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1852 篇

2510.01963 2025-10-03 cs.SD cs.LG 57%

Bias beyond Borders: Global Inequalities in AI-Generated Music

Ahmet Solak, Florian Grötschla, Luca A. Lanzendörfer, Roger Wattenhofer

机构 * ETH Zurich(苏黎世联邦理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12344 2025-10-02 cs.LG cs.DC 57%

CYCle: Choosing Your Collaborators Wisely to Enhance Collaborative Fairness in Decentralized Learning

Nurbek Tastan, Samuel Horvath, Karthik Nandakumar

机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(莫德尔·本·扎耶德人工智能大学) Michigan State University (MSU)(密歇根州立大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments Published in TMLR 08/2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13758 2025-09-30 cs.CY 57%

Towards Evaluting Fake Reasoning Bias in Language Models

Qian Wang, Zhenheng Tang, Zhanzhi Lou, Nuo Chen, Wenxuan Wang, Bingsheng He

专题命中 AI治理与伦理 :DPO(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21542 2025-09-29 cs.HC cs.AI 57%

Psychological and behavioural responses in human-agent vs. human-human interactions: a systematic review and meta-analysis

Jianan Zhou, Fleur Corbett, Joori Byun, Talya Porat, Nejra van Zalk

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21207 2025-09-26 cs.LG 57%

From Physics to Machine Learning and Back: Part II - Learning and Observational Bias in PHM

Olga Fink, Ismail Nejjar, Vinay Sharma, Keivan Faghih Niresi, Han Sun, Hao Dong, Chenghao Xu, Amaury Wei, Arthur Bizzi, Raffael Theiler, Yuan Tian, Leandro Von Krannichfeldt, Zhan Ma, Sergei Garmaev, Zepeng Zhang, Mengjie Zhao

机构 * Intelligent Maintenance and Operations Systems Lab, EPFL, Lausanne, Switzerland(智能维护与运营系统实验室,EPFL,拉沃斯纳,瑞士)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16151 2025-09-22 cs.LG cs.CR 57%

Automated Cyber Defense with Generalizable Graph-based Reinforcement Learning Agents

Isaiah J. King, Benjamin Bowman, H. Howie Huang

机构 * Cybermonic LLC(Cybermonic公司) The George Washington University(乔治华盛顿大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15803 2025-09-22 cs.CV cs.AI 57%

CIDER: A Causal Cure for Brand-Obsessed Text-to-Image Models

Fangjian Shen, Zifeng Liang, Chao Wang, Wushao Wen

机构 * School of Computer Science(计算机科学学院) Engineering, Sun Yat-sen University, Guangzhou, China(工程学院,中山大学,广州,中国)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 5 pages, 7 figures, submitted to ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12566 2025-09-22 cs.AI 57%

Exploring the Impact of Personality Traits on LLM Bias and Toxicity

Shuo Wang, Renhao Li, Xi Chen, Yulin Yuan, Derek F. Wong, Min Yang

机构 * University of Macau(澳门大学) Shenzhen Key Laboratory for High Performance Data Mining, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳高性能数据挖掘重点实验室,深圳先进技术研究院,中国科学院) Nanyang Technological University(南洋理工大学) Department of Chinese Language and Literature, University of Macau(中文语言文学系,澳门大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12104 2025-09-16 cs.AI 57%

JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference

Zongyue Xue, Siyuan Zheng, Shaochun Wang, Yiran Hu, Shenran Wang, Yuxin Yao, Haitao Li, Qingyao Ai, Yiqun Liu, Yun Liu, Weixing Shen

机构 * Tsinghua University(清华大学) Yale Law School(耶鲁法学院) Shanghai Jiaotong University(上海交通大学) University of Waterloo(滑铁卢大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments This paper has been accepted at CIKM 2025 (Demo Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05929 2025-09-11 cs.CY 57%

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support

Keyang Qian, Shiqi Liu, Tongguang Li, Mladen Raković, Xinyu Li, Rui Guan, Inge Molenaar, Sadia Nawaz, Zachari Swiecki, Lixiang Yan, Dragan Gašević

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Journal ref Computers & Education, Volume 240, 2026, 105448

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10160 2025-09-11 cs.CV cs.AI 57%

PriorCLIP: Visual Prior Guided Vision-Language Model for Remote Sensing Image-Text Retrieval

Jiancheng Pan, Muyuan Ma, Qing Ma, Cong Bai, Shengyong Chen

机构 * IEEE Publication Technology Department(IEEE出版技术部)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04993 2025-09-08 cs.MA cs.AI 57%

LLM Enabled Multi-Agent System for 6G Networks: Framework and Method of Dual-Loop Edge-Terminal Collaboration

Zheyan Qu, Wenbo Wang, Zitong Yu, Boquan Sun, Yang Li, Xing Zhang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments This paper has been accepted by IEEE Communications Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20201 2025-09-08 cs.CL 57%

Social Bias in Multilingual Language Models: A Survey

Lance Calvin Lim Gamboa, Yue Feng, Mark Lee

机构 * School of Computer Science, University of Birmingham(伯明翰大学计算机科学学院) Department of Information Systems and Computer Science, Ateneo de Manila University(马尼拉大学信息系统与计算机科学系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted into EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02025 2025-09-05 cs.DC cs.AI 57%

Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling

Prachi Jadhav, Hongwei Jin, Ewa Deelman, Prasanna Balaprakash

机构 * University of Tennessee, Knoxville\ Ridge National Laboratory Oak Ridge, TN USA Argonne National Laboratory Lemont, IL USA University of Southern California Los Angeles, CA USA Oak Ridge National Laboratory Oak Ridge, TN USA University of Tennessee, Knoxville\ Ridge National Laboratory Argonne National Laboratory University of Southern California Oak Ridge National Laboratory

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 10 pages, 6 figures, work under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02007 2025-09-03 cs.AI 57%

mFARM: Towards Multi-Faceted Fairness Assessment based on HARMs in Clinical Decision Support

Shreyash Adappanavar, Krithi Shailya, Gokul S Krishnan, Sriraam Natarajan, Balaraman Ravindran

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02080 2025-09-01 eess.AS cs.AI 57%

Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological Knowledge

Aditya Kamlesh Parikh, Cristian Tejedor-Garcia, Catia Cucchiarini, Helmer Strik

机构 * Centre for Language Studies(语言研究中心)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments Accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)

Journal ref https://www.isca-archive.org/interspeech_2025/parikh25_interspeech.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18765 2025-08-28 cs.LG 57%

Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement

Suyash Gaurav, Jukka Heikkonen, Jatin Chaudhary

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13042 2025-08-27 cs.CY 57%

How Do AI Companies "Fine-Tune" Policy? Examining Regulatory Capture in AI Governance

Kevin Wei, Carson Ezell, Nick Gabrieli, Chinmay Deshpande

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 39 pages (14 pages main text), 3 figures, 9 tables. To be published in the Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, & Society (AIES)

Journal ref Proc. AAAI/ACM Conf. AI, Ethics & Soc., 7 (2024) 1539-1555

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16642 2025-08-26 cs.CY 57%

AI as IA: The use and abuse of artificial intelligence (AI) for human enhancement through intellectual augmentation (IA)

Alexandre Erler, Vincent C. Müller

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Journal ref (2023) in Marcello Ienca and Fabrice Jotterand (eds.), The Routledge Handbook of the Ethics of Human Enhancement (London: Routledge), 187-99

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16013 2025-08-25 cs.CL 57%

Political Ideology Shifts in Large Language Models

Pietro Bernardelle, Stefano Civelli, Leon Fröhling, Riccardo Lunardi, Kevin Roitero, Gianluca Demartini

机构 * The University of Queensland(昆士兰大学) University of Udine(乌迪内大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14415 2025-08-21 cs.AI 57%

The Agent Behavior: Model, Governance and Challenges in the AI Digital Age

Qiang Zhang, Pei Yan, Yijia Xu, Chuanpo Fu, Yong Fang, Yang Liu

机构 * School of Cyber Science and Engineering, Sichuan University, China(计算机科学与工程学院,四川大学) College of Computing and Data Science, Nanyang Technological University, Sinapore(计算与数据科学学院,南洋理工大学) College of Electronics and Information Engineering, Shenzhen University, China(电子与信息工程学院,深圳大学) Department of Computer Science and Technology, Tsinghua University, China(计算机科学与技术系,清华大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13743 2025-08-20 cs.CL 57%

Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA

Kaiwei Zhang, Qi Jia, Zijian Chen, Wei Sun, Xiangyang Zhu, Chunyi Li, Dandan Zhu, Guangtao Zhai

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12174 2025-08-19 cs.CY 57%

Urban AI Governance Must Embed Legal Reasonableness for Democratic and Sustainable Cities

Rashid Mushkani

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11824 2025-08-19 cs.SE cs.AI cs.CR cs.PF 57%

Rethinking Autonomy: Preventing Failures in AI-Driven Software Engineering

Satyam Kumar Navneet, Joydeep Chandra

机构 * Department of CSE Chandigarh University Mohali, India(计算机科学与工程系 奇纳格里大学 莫哈利,印度) Department of CST Tsinghua University Beijing, China(计算机科学与技术系 清华大学 北京,中国)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11262 2025-08-18 cs.CV cs.AI 57%

Vision-Language Models display a strong gender bias

Aiswarya Konavoor, Raj Abhijit Dandekar, Rajat Dandekar, Sreedath Panat

机构 * Togo AI Labs(Togo人工智能实验室) Vizuara AI Labs(Vizuara人工智能实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10007 2025-08-15 cs.CL stat.ME 57%

Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models

Y. Lyu, D. Combs, D. Neumann, Y. C. Leong

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments We have no known conflict of interest

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09019 2025-08-13 cs.AI 57%

Activation Steering for Bias Mitigation: An Interpretable Approach to Safer LLMs

Shivam Dubey

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯学院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08262 2025-08-13 cs.CL 57%

Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models

Alaa Alhamzeh, Mays Al Rebdawi

机构 * First Author Affiliation(第一作者机构) Second Author Affiliation(第二作者机构)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments 8 pages, 4 figures, Passau uni, Master thesis in NLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16170 2025-08-12 cs.AI 57%

Learning How to Vote with Principles: Axiomatic Insights Into the Collective Decisions of Neural Networks

Levin Hornischer, Zoi Terzopoulou

机构 * Munich Center for Mathematical Philosophy, LMU Munich Munich Germany GATE, CNRS, Universit\'e Jean Monnet, Universit\'e Lumiere Lyon 2 Saint-Etienne France Munich Center for Mathematical Philosophy, LMU Munich GATE, CNRS, Universit\'e Jean Monnet, Universit\'e Lumiere Lyon 2

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 44 pages, 21 figures, 14 tables. Updated and published version

Journal ref Journal of Artificial Intelligence Research 83, Article 25 (August 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06479 2025-08-11 cs.CY 57%

The Problem of Atypicality in LLM-Powered Psychiatry

Bosco Garcia, Eugene Y. S. Chua, Harman Singh Brah

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Preprint of 8/8/2025 -- please cite published version. This article has been published in the Journal of Medical Ethics (2025) following peer review and can also be viewed on the journal's website at 10.1136/jme-2025-110972

详情

展开后加载摘要…

URL PDF HTML 收藏