arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-16 至 2025-09-16 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 5 篇

2509.11620 2025-09-16 cs.CL cs.CY 81%

AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment

Kun Li, Lai-Man Po, Hongzheng Yang, Xuyuan Xu, Kangcheng Liu, Yuzhi Zhao

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.CY

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11648 2025-09-16 cs.CL cs.AI cs.CY 67%

EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI

Sai Kartheek Reddy Kasu

机构 * IIIT Dharwad(德瓦德理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17044 2025-09-16 cs.CY cs.AI cs.LG 67%

Approaches to Responsible Governance of GenAI in Organizations

Dhari Gandhi, Himanshu Joshi, Lucas Hartman, Shabnam Hassani

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) Western University(西部大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10653 2025-09-16 cs.CY cs.AI 62%

SCOR: A Framework for Responsible AI Innovation in Digital Ecosystems

Mohammad Saleh Torkestani, Taha Mansouri

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Proceeding of The British Academy of Management Conference 2025, University of Kent, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12104 2025-09-16 cs.AI 57%

JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference

Zongyue Xue, Siyuan Zheng, Shaochun Wang, Yiran Hu, Shenran Wang, Yuxin Yao, Haitao Li, Qingyao Ai, Yiqun Liu, Yun Liu, Weixing Shen

机构 * Tsinghua University(清华大学) Yale Law School(耶鲁法学院) Shanghai Jiaotong University(上海交通大学) University of Waterloo(滑铁卢大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments This paper has been accepted at CIKM 2025 (Demo Track)

详情

展开后加载摘要…

URL PDF HTML 收藏