arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-07 至 2025-10-07 共收录 9 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 9 篇

2510.04528 2025-10-07 cs.CR cs.AI 79%

Unified Threat Detection and Mitigation Framework (UTDMF): Combating Prompt Injection, Deception, and Bias in Enterprise-Scale Transformers

Santhosh KumarRavindran

机构 * Microsoft Corporation(微软公司)

专题命中 AI治理与伦理 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04073 2025-10-07 cs.AI 79%

Moral Anchor System: A Predictive Framework for AI Value Alignment and Drift Prevention

Santhosh Kumar Ravindran

机构 * Microsoft Corporation(微软公司)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

Comments 11 pages Includes simulations with over 4 million steps

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06303 2025-10-07 cs.CY cs.AI cs.CL cs.LG 70%

On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions

Dang Nguyen, Chenhao Tan

机构 * Department of Computer Science University of Chicago(计算机科学系芝加哥大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 21 pages, 15 figures, 14 tables. Accepted as a conference paper at COLM 2025. Camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10127 2025-10-07 cs.CL cs.AI cs.LG 67%

Population-Aligned Persona Generation for LLM-based Social Simulation

Zhengyu Hu, Jianxun Lian, Zheyuan Xiao, Max Xiong, Yuxuan Lei, Tianfu Wang, Kaize Ding, Ziang Xiao, Nicholas Jing Yuan, Xing Xie

机构 * HKUST(香港科技大学) Microsoft Research Asia(微软亚洲研究院) Duke University(杜克大学) Northwestern University(西北大学) Johns Hopkins University(约翰霍普金斯大学) Microsoft(微软)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18562 2025-10-07 cs.CL cs.AI 66%

From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test

Xunlian Dai, Li Zhou, Benyou Wang, Haizhou Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院)

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CL、cs.AI

Comments Cultural Analysis, Cultural Alignment, Word Association Test, Large Language Models. Accepted by EMNLP 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10659 2025-10-07 cs.SI cs.AI cs.CL cs.MA 62%

Network Formation and Dynamics Among Multi-LLMs

Marios Papachristou, Yuan Yuan

机构 * Department of Information Systems, W.P. Carey School of Business, Arizona State University, Tempe, AZ, USA(亚利桑那州立大学信息系统系,W.P. Carey商学院,Tempe分校) Department of Computer Science, Cornell University, Ithaca, NY, USA(康奈尔大学计算机科学系) Graduate School of Management, University of California Davis, Davis, CA, USA(加州大学戴维斯分校管理研究生院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at PNAS Nexus

Journal ref PNAS Nexus 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03368 2025-10-07 cs.CY cs.AI 62%

An Adaptive Responsible AI Governance Framework for Decentralized Organizations

Kiana Jafari Meimandi, Anka Reuel, Gabriela Aranguiz-Dias, Hatim Rahama, Ala-Eddine Ayadi, Xavier Boullier, Jérémy Verdo, Louis Montanie, Mykel Kochenderfer

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04577 2025-10-07 cs.SD cs.LG cs.MM eess.AS 57%

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

Juncheng Wang, Chao Xu, Cheng Yu, Zhe Hu, Haoyu Xie, Guoqi Yu, Lei Shang, Shujun Wang

机构 * The Hong Kong Polytechnic University(香港理工大学) Alibaba Group(阿里巴巴集团)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04038 2025-10-07 eess.SY cs.SY 50%

Distributed MPC-based Coordination of Traffic Perimeter and Signal Control: A Lexicographic Optimization Approach

Viet Hoang Pham, Hyo-Sung Ahn

专题命中 AI治理与伦理 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏