arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2508.05913 2025-08-11 cs.HC cs.AI cs.CL 62%

Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction

Stefan Pasch, Min Chul Cha

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03970 2025-08-07 cs.CL cs.AI 62%

Data and AI governance: Promoting equity, ethics, and fairness in large language models

Alok Abhishek, Lisa Erickson, Tushar Bandopadhyay

机构 * Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments Published in MIT Science Policy Review 6, 139-146 (2025)

Journal ref MIT Science Policy Review, 6. (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03292 2025-08-06 cs.CL cs.AI 62%

Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes

Shahed Masoudian, Gustavo Escobedo, Hannah Strauss, Markus Schedl

机构 * Johannes Kepler University (JKU)(约翰内斯·开普勒大学) Linz Institute of Technology (LIT)(林茨技术研究所) University of Innsbruck(因斯布鲁克大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09053 2025-08-06 cs.AI cs.GT cs.LG 62%

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers

Haoran Sun, Yusen Wu, Peng Wang, Wei Chen, Yukun Cheng, Xiaotie Deng, Xu Chu

机构 * CFCS, School of Computer Science, Peking University(计算机科学系,北京大学) School of Business, Jiangnan University(商学院,江南大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments A shorter conference version is published in IJCAI 2025, titled 'Game Theory Meets Large Language Models: A Systematic Survey'

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06072 2025-08-04 cs.CL cs.AI 62%

A Survey on Post-training of Large Language Models

Guiyao Tie, Zeli Zhao, Dingjie Song, Fuyang Wei, Rong Zhou, Yurou Dai, Wen Yin, Zhejian Yang, Jiangyue Yan, Yao Su, Zhenhan Dai, Yifeng Xie, Yihan Cao, Lichao Sun, Pan Zhou, Lifang He, Hechang Chen, Yu Zhang, Qingsong Wen, Tianming Liu, Neil Zhenqiang Gong, Jiliang Tang, Caiming Xiong, Heng Ji, Philip S. Yu, Jianfeng Gao

机构 * Huazhong University of Science and Technology(华中科技大学) Lehigh University(莱斯大学) The University of Hong Kong(香港大学) Jilin University(吉林大学) Southern University of Science and Technology(南方科技大学) Worcester Polytechnic Institute(沃思堡理工学院) LinkedIn Corporation(领英公司) Squirrel Ai Learning University of Georgia(佐治亚大学) Duke University(杜克大学) Michigan State University(密歇根州立大学) Salesforce Research(Salesforce研究) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Illinois at Chicago(伊利诺伊大学芝加哥分校) Microsoft Research(微软研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 87 pages, 21 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22445 2025-07-31 cs.CL cs.AI 62%

AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by gpt-4o-mini

Jill Walker Rettberg, Hermann Wigers

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments This project has received funding from the European Union's Horizon 2020 research and innovation programme under grant agreement number 101142306. The project is also supported by the Center for Digital Narrative, which is funded by the Research Council of Norway through its Centres of Excellence scheme, project number 332643

Journal ref Open Research Europe 2025, 5:202 [version 1; peer review: awaiting peer review]

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21083 2025-07-30 cs.CL cs.AI 62%

ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs

Franck Bardol

机构 * Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17788 2025-07-25 cs.LG cs.AI 62%

Adaptive Repetition for Mitigating Position Bias in LLM-Based Ranking

Ali Vardasbi, Gustavo Penha, Claudia Hauff, Hugues Bouchard

机构 * Spotify

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17787 2025-07-25 cs.LG cs.AI 62%

Hyperbolic Deep Learning for Foundation Models: A Survey

Neil He, Hiren Madhu, Ngoc Bui, Menglin Yang, Rex Ying

机构 * Yale University(耶鲁大学) Hong Kong University of Science(香港科学大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments 11 Pages, SIGKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15901 2025-07-23 cs.AI cs.CY cs.MA 62%

Advancing Responsible Innovation in Agentic AI: A study of Ethical Frameworks for Household Automation

Joydeep Chandra, Satyam Kumar Navneet

机构 * Department of CST Tsinghua University Beijing, China(计算机科学与技术系 清华大学 北京中国) Department of CSE Chandigarh University Mohali, India(计算机科学与工程系 印度昌迪加尔大学 摩哈利)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15885 2025-07-23 cs.AI cs.HC cs.LG 62%

ADEPTS: A Capability Framework for Human-Centered Agent Design

Pierluca D'Oro, Caley Drooff, Joy Chen, Joseph Tighe

机构 * FAIR at Meta(Meta 的 FAIR)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11415 2025-07-15 cs.CL cs.AI 62%

Political Bias in LLMs: Unaligned Moral Values in Agent-centric Simulations

Simon Münker

机构 * Journal for Language Technology and Computational Linguistics(语言技术与计算语言学期刊)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 2 tables

Journal ref JLCL 2025, Band 38(2)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08908 2025-07-15 cs.CY cs.ET cs.LG 62%

The Engineer's Dilemma: A Review of Establishing a Legal Framework for Integrating Machine Learning in Construction by Navigating Precedents and Industry Expectations

M. Z. Naser

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05443 2025-07-03 cs.CL cs.AI 62%

A survey of textual cyber abuse detection using cutting-edge language models and large language models

Jose A. Diaz-Garcia, Joao Paulo Carvalho

机构 * Department of Computer Science and A.I, University of Granada(计算机科学与人工智能系,格拉纳达大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments 37 pages, under review in WIREs Data Mining and Knowledge Discovery

Journal ref Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery (2025), 15(3), e70029

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15620 2025-06-19 cs.LG cs.AI 62%

GFLC: Graph-based Fairness-aware Label Correction for Fair Classification

Modar Sulaiman, Kallol Roy

机构 * University of Tartu, Institute of Computer Science(塔尔图大学计算机科学研究所)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 25 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14645 2025-06-18 cs.CL cs.CY 62%

Passing the Turing Test in Political Discourse: Fine-Tuning LLMs to Mimic Polarized Social Media Comments

. Pazzaglia, V. Vendetti, L. D. Comencini, F. Deriu, V. Modugno

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11627 2025-06-16 cs.CV cs.AI cs.LG 62%

Evaluating Fairness and Mitigating Bias in Machine Learning: A Novel Technique using Tensor Data and Bayesian Regression

Kuniko Paxton, Koorosh Aslansefat, Dhavalkumar Thakker, Yiannis Papadopoulos

机构 * School of Computer Science and DAIM University of Hull(计算机科学学院和DAIM赫尔大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02221 2025-06-12 cs.LG cs.AI stat.ML 62%

Bias Detection via Maximum Subgroup Discrepancy

Jiří Němeček, Mark Kozdoba, Illia Kryvoviaz, Tomáš Pevný, Jakub Mareček

机构 * Czech Technical University in Prague, Faculty of Electrical Engineering(捷克技术大学布拉格电子工程学院)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 6 figures

Journal ref Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08231 2025-06-11 cs.LG cs.AI cs.PF 62%

Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework

Melissa Estevez, Nisha Singh, Lauren Dyson, Blythe Adamson, Qianyu Yuan, Megan W. Hildner, Erin Fidyk, Olive Mbah, Farhad Khan, Kathi Seidl-Rathkopf, Aaron B. Cohen

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 3 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05728 2025-06-04 cs.CY cs.AI 62%

Political Neutrality in AI Is Impossible- But Here Is How to Approximate It

Jillian Fisher, Ruth E. Appel, Chan Young Park, Yujin Potter, Liwei Jiang, Taylor Sorensen, Shangbin Feng, Yulia Tsvetkov, Margaret E. Roberts, Jennifer Pan, Dawn Song, Yejin Choi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments Code: https://github.com/jfisher52/Approximation_Political_Neutrality

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19331 2025-06-02 cs.LG cs.CY 62%

Friends in Unexpected Places: Enhancing Local Fairness in Federated Learning through Clustering

Yifan Yang, Ali Payani, Parinaz Naghizadeh

机构 * The Ohio State University(俄亥俄州立大学) Cisco Research(思科研究) UC, San Diego(圣地亚哥大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21570 2025-05-29 cs.CY cs.AI 62%

Beyond Explainability: The Case for AI Validation

Dalit Ken-Dror Feldman, Daniel Benoliel

机构 * University of Haifa(海法大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21562 2025-05-29 cs.CY cs.AI cs.HC 62%

Enhancing Selection of Climate Tech Startups with AI -- A Case Study on Integrating Human and AI Evaluations in the ClimaTech Great Global Innovation Challenge

Jennifer Turliuk, Alejandro Sevilla, Daniela Gorza, Tod Hynes

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21537 2025-05-29 cs.CY cs.AI 62%

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models

Hao Sun, Yunyi Shen, Mihaela van der Schaar

机构 * University of Cambridge(剑桥大学) MIT(麻省理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20692 2025-05-28 cs.HC cs.AI cs.CL 62%

Can we Debias Social Stereotypes in AI-Generated Images? Examining Text-to-Image Outputs and User Perceptions

Saharsh Barve, Andy Mao, Jiayue Melissa Shi, Prerna Juneja, Koustuv Saha

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15481 2025-05-27 cs.CL cs.CY 62%

Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency

Yiran Liu, Ke Yang, Zehan Qi, Xiao Liu, Yang Yu, ChengXiang Zhai

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19776 2025-05-27 cs.CL cs.AI 62%

Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification

Akram Elbouanani, Evan Dufraisse, Adrian Popescu

机构 * Université Paris-Saclay, CEA, List(巴黎-萨克雷大学、欧洲原子能机构、List)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments To be published in the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18494 2025-05-27 cs.LG cs.AI 62%

FedHL: Federated Learning for Heterogeneous Low-Rank Adaptation via Unbiased Aggregation

Zihao Peng, Jiandian Zeng, Boyuan Li, Guo Li, Shengbo Chen, Tian Wang

机构 * Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai(人工智能与未来网络研究院,北京师范大学,珠海) School of Computer Science and Artificial Intelligence, Zhengzhou University(计算机科学与人工智能学院,郑州大学) School of Software, Nanchang University(软件学院,南昌大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06790 2025-05-23 cs.CY cs.CL cs.HC 62%

Large-scale moral machine experiment on large language models

Muhammad Shahrul Zaim bin Ahmad, Kazuhiro Takemoto

机构 * Department of Bioscience and Bioinformatics, Kyushu Institute of Technology(九州工科大学生物科学与生物信息学系) Faculty of Engineering and Technology, Multimedia University(多媒体大学工程与技术学院) Data Science and AI Research Center, Kyushu Institute of Technology(九州工科大学数据科学与人工智能研究中心)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

Comments 21 pages, 6 figures

Journal ref PLoS One 20, e0322776 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15031 2025-05-22 cs.CL cs.AI cs.HC cs.IR 62%

Are the confidence scores of reviewers consistent with the review content? Evidence from top conference proceedings in AI

Wenqing Wu, Haixu Xi, Chengzhi Zhang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref Scientometrics, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏