arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-11 至 2025-11-11 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 7 篇

2511.06023 2025-11-11 cs.CL 77%

Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data

Deng Yixuan, Ji Xiaoqiang

机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)科学与工程学院) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院) Shenzhen Institute of Artificial Intelligence and Robotics for Society, China(深圳人工智能与机器人研究院)

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04698 2025-11-11 cs.CL cs.AI 62%

multiMentalRoBERTa: A Fine-tuned Multiclass Classifier for Mental Health Disorder

K M Sajjadul Islam, John Fields, Praveen Madiraju

机构 * Marquette University, WI, USA(马quette大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted in IEEE Big Data, 8-11 December, 2025 @ Macau SAR, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03262 2025-11-11 cs.CL cs.LG 62%

REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Jian Hu, Jason Klein Liu, Haotian Xu, Wei Shen

专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.CL、cs.LG

Comments refactor

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05766 2025-11-11 cs.AI cs.CL econ.GN q-fin.EC 62%

Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs

Felipe Valencia-Clavijo

机构 * Dataplicada

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23628 2025-11-11 physics.soc-ph cs.CY cs.DS cs.GT econ.TH 57%

Matchings Under Biased and Correlated Evaluations

Amit Kumar, Nisheeth K. Vishnoi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments To appear in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13182 2025-11-11 cs.LO cs.AI 57%

Information Science Principles of Machine Learning: A Causal Chain Meta-Framework Based on Formalized Information Mapping

Jianfeng Xu

机构 * Koguan School of Law, China Institute for Smart Justice, School of Computer Science, Shanghai Jiao Tong University(柯 guar 法学院、中国智能正义研究院、计算机科学学院、上海交通大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15857 2025-11-11 cs.SI 50%

Investigating Prosocial Behavior Theory in LLM Agents under Policy-Induced Inequities

Yujia Zhou, Hexi Wang, Qingyao Ai, Zhen Wu, Yiqun Liu

专题命中 AI治理与伦理 :alignment(abstract)

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏