arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-17 至 2025-10-17 共收录 9 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 9 篇

2509.24065 2025-10-17 cs.CY 90%

AI Safety, Alignment, and Ethics (AI SAE)

Dylan Waldner

专题命中 AI治理与伦理 :alignment(title,abstract);safety(title);AI safety(title);分类 cs.CY

Comments 46 pages, 9 figures (including appendix), 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13931 2025-10-17 cs.CL 79%

Robust or Suggestible? Exploring Non-Clinical Induction in LLM Drug-Safety Decisions

Siying Liu, Shisheng Zhang, Indu Bala

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CL

Comments Preprint of a paper accepted as a poster at the NeurIPS 2025 Workshop on Generative AI for Health (GenAI4Health). The final camera-ready workshop version may differ. Licensed under CC BY 4.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07887 2025-10-17 cs.CL cs.AI 73%

Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge

Riccardo Cantini, Alessio Orsino, Massimo Ruggiero, Domenico Talia

机构 * University of Calabria(卡利博利亚大学)

专题命中 AI治理与伦理 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Journal ref Cantini, R., Orsino, A., Ruggiero, M., Talia, D. Benchmarking adversarial robustness to bias elicitation in large language models: scalable automated assessment with LLM-as-a-judge. Mach Learn 114, 249 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14053 2025-10-17 cs.AI 70%

Position: Require Frontier AI Labs To Release Small "Analog" Models

Shriyash Upadhyay, Chaithanya Bandi, Narmeen Oozeer, Philip Quirke

机构 * Frontier AI Labs(前沿人工智能实验室)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14106 2025-10-17 cs.AI cs.CL cs.GT 62%

Generating Fair Consensus Statements with Social Choice on Token-Level MDPs

Carter Blair, Kate Larson

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08236 2025-10-17 cs.LG cs.AI 62%

The Hidden Bias: A Study on Explicit and Implicit Political Stereotypes in Large Language Models

Konrad Löhr, Shuzhou Yuan, Michael Färber

机构 * Technische Universität Dresden(德累斯顿技术大学) Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI)(可扩展数据与人工智能研究中心(ScaDS.AI))

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14443 2025-10-17 cs.SD cs.AI eess.AS 57%

Big Data Approaches to Bovine Bioacoustics: A FAIR-Compliant Dataset and Scalable ML Framework for Precision Livestock Welfare

Mayuri Kate, Suresh Neethirajan

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 40 pages, 14 figures, 9 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10623 2025-10-17 cs.LG cs.CV 57%

Flows and Diffusions on the Neural Manifold

Daniel Saragih, Deyu Cao, Tejas Balaji

机构 * Queen’s University and Vector Institute(女王大学和向量研究所) University of Tokyo(东京大学) University of Toronto(多伦多大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 43 pages, 11 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13699 2025-10-17 cs.IR cs.AI 57%

A Comprehensive Review of Recommender Systems: Transitioning from Theory to Practice

Shaina Raza, Mizanur Rahman, Safiullah Kamawal, Armin Toroghi, Ananya Raval, Farshad Navah, Amirmohammad Kazemeini

机构 * Vector Institute(向量研究所) Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments we quarterly update of this literature

详情

展开后加载摘要…

URL PDF HTML 收藏