arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-26 至 2025-09-26 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 3 篇

2509.20394 2025-09-26 cs.CY cs.AI cs.CL cs.CR 75%

Blueprints of Trust: AI System Cards for End to End Transparency and Governance

Huzaifa Sidhpurwala, Emily Fox, Garth Mollett, Florencio Cano Gabarda, Roman Zhukov

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21075 2025-09-26 cs.CY cs.AI cs.CL cs.DC cs.HC cs.LG 70%

Communication Bias in Large Language Models: A Regulatory Perspective

Adrian Kuenzler, Stefan Schmid

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21207 2025-09-26 cs.LG 57%

From Physics to Machine Learning and Back: Part II - Learning and Observational Bias in PHM

Olga Fink, Ismail Nejjar, Vinay Sharma, Keivan Faghih Niresi, Han Sun, Hao Dong, Chenghao Xu, Amaury Wei, Arthur Bizzi, Raffael Theiler, Yuan Tian, Leandro Von Krannichfeldt, Zhan Ma, Sergei Garmaev, Zepeng Zhang, Mengjie Zhao

机构 * Intelligent Maintenance and Operations Systems Lab, EPFL, Lausanne, Switzerland(智能维护与运营系统实验室,EPFL,拉沃斯纳,瑞士)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏