arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-27 至 2025-08-27 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 3 篇

2508.18886 2025-08-27 cs.CV 78%

Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models

Yuexuan Xia, Benteng Ma, Jiang He, Zhiyong Wang, Qi Dou, Yong Xia

机构 * National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室) Northwestern Polytechnical University(西北工业大学) Huiying Medical Technology Company Ltd.(慧影医疗技术有限公司) The School of Computer Science(计算机学院) The University of Sydney(悉尼大学) Department of Computer Science and Engineering(计算机科学与工程系) The Chinese University of Hong Kong(香港中文大学)

专题命中 AI治理与伦理 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13042 2025-08-27 cs.CY 57%

How Do AI Companies "Fine-Tune" Policy? Examining Regulatory Capture in AI Governance

Kevin Wei, Carson Ezell, Nick Gabrieli, Chinmay Deshpande

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 39 pages (14 pages main text), 3 figures, 9 tables. To be published in the Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, & Society (AIES)

Journal ref Proc. AAAI/ACM Conf. AI, Ethics & Soc., 7 (2024) 1539-1555

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18600 2025-08-27 cs.GT cs.MA econ.GN q-fin.EC 50%

Bias-Adjusted LLM Agents for Human-Like Decision-Making via Behavioral Economics

Ayato Kitadai, Yusuke Fukasawa, Nariaki Nishino

专题命中 AI治理与伦理 :alignment(abstract)

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏