arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2411.13584 2024-11-22 cs.CL cs.AI 62%

AddrLLM: Address Rewriting via Large Language Model on Nationwide Logistics Data

Qinchen Yang, Zhiqing Hong, Dongjiang Cao, Haotian Wang, Zejun Xie, Tian He, Yunhuai Liu, Yu Yang, Desheng Zhang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by KDD'25 ADS Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12820 2024-11-21 cs.AI cs.CY 62%

Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation

Peter Barnett, Lisa Thiergart

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09788 2024-11-18 cs.HC cs.AI cs.CY 62%

AI-Driven Human-Autonomy Teaming in Tactical Operations: Proposed Framework, Challenges, and Future Directions

Desta Haileselassie Hagos, Hassan El Alami, Danda B. Rawat

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments Submitted for review to the Proceedings of the IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07398 2024-11-13 cs.CL cs.AI cs.SE 62%

Beyond Keywords: A Context-based Hybrid Approach to Mining Ethical Concern-related App Reviews

Aakash Sorathiya, Gouri Ginde

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12789 2024-11-11 cs.LG cs.AI 62%

Fairness Without Harm: An Influence-Guided Active Sampling Approach

Jinlong Pang, Jialu Wang, Zhaowei Zhu, Yuanshun Yao, Chen Qian, Yang Liu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14106 2024-11-11 cs.AI cs.LG 62%

Learning Human-like Representations to Enable Learning Human Values

Andrea Wynn, Ilia Sucholutsky, Thomas L. Griffiths

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03651 2024-11-07 cs.AI cs.GT cs.LG 62%

Policy Aggregation

Parand A. Alamdari, Soroush Ebadian, Ariel D. Procaccia

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03284 2024-11-06 cs.AI cs.CL cs.MA 62%

SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents

Dawei Li, Zhen Tan, Peijia Qian, Yifan Li, Kumar Satvik Chaudhary, Lijie Hu, Jiayi Shen

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13517 2024-11-06 cs.CL cs.AI 62%

Bias in the Mirror: Are LLMs opinions robust to their own adversarial attacks ?

Virgile Rennard, Christos Xypolopoulos, Michalis Vazirgiannis

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21279 2024-10-30 cs.CY cs.AI 62%

Comparative Global AI Regulation: Policy Perspectives from the EU, China, and the US

Jon Chun, Christian Schroeder de Witt, Katherine Elkins

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 36 pages, 11 figures and tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10284 2024-10-22 cs.CY cs.AI cs.RO 62%

Trust or Bust: Ensuring Trustworthiness in Autonomous Weapon Systems

Kasper Cools, Clara Maathuis

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments Accepted as a workshop paper at MILCOM 2024, 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13250 2024-10-18 cs.HC cs.AI cs.CY 62%

Perceptions of Discriminatory Decisions of Artificial Intelligence: Unpacking the Role of Individual Characteristics

Soojong Kim

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13138 2024-10-18 cs.CL cs.CR cs.CY 62%

Data Defenses Against Large Language Models

William Agnew, Harry H. Jiang, Cella Sum, Maarten Sap, Sauvik Das

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12848 2024-10-18 cs.CL cs.AI cs.HC 62%

Prompt Engineering a Schizophrenia Chatbot: Utilizing a Multi-Agent Approach for Enhanced Compliance with Prompt Instructions

Per Niklas Waaler, Musarrat Hussain, Igor Molchanov, Lars Ailo Bongo, Brita Elvevåg

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08760 2024-10-16 cs.CL cs.AI 62%

The Generation Gap: Exploring Age Bias in the Value Systems of Large Language Models

Siyang Liu, Trish Maturi, Bowen Yi, Siqi Shen, Rada Mihalcea

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 5 pages

Journal ref The 2024 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12232 2024-10-08 cs.AI cs.CL 62%

"You Gotta be a Doctor, Lin": An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations

Huy Nghiem, John Prindle, Jieyu Zhao, Hal Daumé

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024, 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01819 2024-10-04 cs.CY cs.AI 62%

Strategic AI Governance: Insights from Leading Nations

Dian W. Tjondronegoro

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 21 pages, 3 Figures, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12926 2024-09-24 cs.LG cs.AI 62%

Trusting Fair Data: Leveraging Quality in Fairness-Driven Data Removal Techniques

Manh Khoi Duong, Stefan Conrad

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments The Version of Record of this contribution is published in Springer LNCS 14912 and is available online at https://doi.org/10.1007/978-3-031-68323-7_33

Journal ref Lecture Notes in Computer Science, Vol. 14912 (2024), pp. 375-380. Springer

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13852 2024-09-24 cs.CL cs.AI 62%

Do language models practice what they preach? Examining language ideologies about gendered language reform encoded in LLMs

Julia Watson, Sophia Lee, Barend Beekhuizen, Suzanne Stevenson

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07335 2024-09-12 cs.AI cs.CL 62%

Explanation, Debate, Align: A Weak-to-Strong Framework for Language Model Generalization

Mehrdad Zakershahrak, Samira Ghodratnama

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07467 2024-09-09 cs.CY cs.AI 62%

AI Ethics Principles in Practice: Perspectives of Designers and Developers

Conrad Sanderson, David Douglas, Qinghua Lu, Emma Schleiger, Jon Whittle, Justine Lacey, Glenn Newnham, Stefan Hajkowicz, Cathy Robinson, David Hansen

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments submitted to IEEE Transactions on Technology & Society

Journal ref IEEE Transactions on Technology and Society, Vol. 4, No. 2, pp. 171-187, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12691 2024-09-04 cs.AI cs.CY 62%

Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?

Shayne Longpre, Robert Mahari, Naana Obeng-Marnu, William Brannon, Tobin South, Katy Gero, Sandy Pentland, Jad Kabbara

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments ICML 2024 camera-ready version (Spotlight paper). 9 pages, 2 tables

Journal ref Proceedings of ICML 2024, in PMLR 235:32711-32725. URL: https://proceedings.mlr.press/v235/longpre24b.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12055 2024-08-23 cs.CL cs.LG 62%

Aligning (Medical) LLMs for (Counterfactual) Fairness

Raphael Poulain, Hamed Fayyaz, Rahmatollah Beheshti

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

Comments arXiv admin note: substantial text overlap with arXiv:2404.15149

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11457 2024-08-22 cs.IR cs.AI cs.CL 62%

Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era

Sunhao Dai, Chen Xu, Shicheng Xu, Liang Pang, Zhenhua Dong, Jun Xu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments KDD 2024 Tutorial&Survey; Tutorial Website: https://llm-ir-bias-fairness.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10608 2024-08-21 cs.CL cs.AI 62%

Promoting Equality in Large Language Models: Identifying and Mitigating the Implicit Bias based on Bayesian Theory

Yongxin Deng, Xihe Qiu, Xiaoyu Tan, Jing Pan, Chen Jue, Zhijun Fang, Yinghui Xu, Wei Chu, Yuan Qi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03907 2024-08-08 cs.CL cs.AI 62%

Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models

Shachi H Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Beckage, Hsuan Su, Hung-yi Lee, Lama Nachman

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 6 pages paper content, 17 pages of appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21281 2024-08-01 cs.CY cs.AI 62%

Unlocking the Potential of Binding Corporate Rules (BCRs) in Health Data Transfers

Marcelo Corrales Compagnucci, Mark Fenwick, Helena Haapio

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16903 2024-07-25 cs.CY cs.AI 62%

US-China perspectives on extreme AI risks and global governance

Akash Wasil, Tim Durgin

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13025 2024-07-22 cs.CY cs.AI 62%

From Principles to Practices: Lessons Learned from Applying Partnership on AI's (PAI) Synthetic Media Framework to 11 Use Cases

Claire R. Leibowicz, Christian H. Cardona

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 18 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12999 2024-07-19 cs.CY cs.AI cs.CR 62%

Securing the Future of GenAI: Policy and Technology

Mihai Christodorescu, Ryan Craven, Soheil Feizi, Neil Gong, Mia Hoffmann, Somesh Jha, Zhengyuan Jiang, Mehrdad Saberi Kamarposhti, John Mitchell, Jessica Newman, Emelia Probasco, Yanjun Qi, Khawaja Shams, Matthew Turek

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏