arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-12 至 2025-11-12 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 6 篇

2511.08567 2025-11-12 cs.LG cs.AI 62%

The Path Not Taken: RLVR Provably Learns Off the Principals

Hanqing Zhu, Zhenyu Zhang, Hanxian Huang, DiJia Su, Zechun Liu, Jiawei Zhao, Igor Fedorov, Hamed Pirsiavash, Zhizhou Sha, Jinwon Lee, David Z. Pan, Zhangyang Wang, Yuandong Tian, Kai Sheng Tai

机构 * Meta AI The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments Preliminary version accepted as a spotlight in NeurIPS 2025 Workshop on Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08082 2025-11-12 cs.AI cs.LG econ.GN q-fin.EC 62%

Prudential Reliability of Large Language Models in Reinsurance: Governance, Assurance, and Capital Efficiency

Stella C. Dong

机构 * Reinsurance Analytics(再保险分析)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments 48 pages, 9 figures, 5 tables. Submitted to the Journal of Risk and Insurance (JRI), November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07803 2025-11-12 cs.CY cs.AI 62%

Judging by the Rules: Compliance-Aligned Framework for Modern Slavery Statement Monitoring

Wenhao Xu, Akshatha Arodi, Jian-Yun Nie, Arsene Fansi Tchango

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments To appear at AAAI-26 (Social Impact Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07941 2025-11-12 cs.CV cs.AI 57%

Libra-MIL: Multimodal Prototypes Stereoscopic Infused with Task-specific Language Priors for Few-shot Whole Slide Image Classification

Zhenfeng Zhuang, Fangyu Zhou, Liansheng Wang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02241 2025-11-12 cs.CL 57%

Isolating Culture Neurons in Multilingual Large Language Models

Danial Namazifard, Lukas Galke Poech

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted at IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14686 2025-11-12 cs.CV 50%

From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition

Chen Cai, Tianyi Liu, Jianjun Gao, Wenyang Liu, Kejun Wu, Ruoyu Wang, Yi Wang, Soo Chin Liew

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Huazhong University of Science and Technology(华中科技大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 AI治理与伦理 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏