arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2001.00088 2022-06-14 cs.CY cs.GT cs.LG 62%

AI for Social Impact: Learning and Planning in the Data-to-Deployment Pipeline

Andrew Perrault, Fei Fang, Arunesh Sinha, Milind Tambe

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

Comments AI Magazine, Winter 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.03220 2022-06-08 cs.CY cs.AI 62%

A Transparency Index Framework for AI in Education

Muhammad Ali Chaudhry, Mutlu Cukurova, Rose Luckin

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 17 pages, 4 Figures, 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10785 2022-05-24 cs.CY cs.AI 62%

Responsible Artificial Intelligence -- from Principles to Practice

Virginia Dignum

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments This paper is a curated version of my keynote at the Web Conference 2022. arXiv admin note: substantial text overlap with arXiv:2202.07446

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04279 2022-05-10 cs.CY cs.AI 62%

Aligned with Whom? Direct and social goals for AI systems

Anton Korinek, Avital Balwit

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Prepared for the Oxford Handbook of AI Governance (23 pages, 2 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.02894 2022-05-09 cs.HC cs.AI cs.CL 62%

Interactive Model Cards: A Human-Centered Approach to Model Documentation

Anamaria Crisan, Margaret Drouhard, Jesse Vig, Nazneen Rajani

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments To appear at ACM FAccT'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.09110 2021-12-07 cs.CY cs.AI 62%

Developing Future Human-Centered Smart Cities: Critical Analysis of Smart City Security, Interpretability, and Ethical Challenges

Kashif Ahmad, Majdi Maabreh, Mohamed Ghaly, Khalil Khan, Junaid Qadir, Ala Al-Fuqaha

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.12146 2021-10-20 cs.CY cs.AI 62%

Avoiding Negative Side Effects due to Incomplete Knowledge of AI Systems

Sandhya Saisubramanian, Shlomo Zilberstein, Ece Kamar

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.02972 2021-10-15 cs.LG cs.CY 62%

Empirical observation of negligible fairness-accuracy trade-offs in machine learning for public policy

Kit T. Rodolfa, Hemank Lamba, Rayid Ghani

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

Comments 40 pages, 4 figures, 2 tables, 7 supplementary figures, 4 supplementary tables; revised to improve clarity and discussion

Journal ref Nat Mach Intell 3, 896-904 (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02203 2021-10-05 cs.CY cs.LG 62%

Accuracy-Efficiency Trade-Offs and Accountability in Distributed ML Systems

A. Feder Cooper, Karen Levy, Christopher De Sa

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

Journal ref Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.02928 2021-09-15 cs.CY cs.AI 62%

Toward a Rational and Ethical Sociotechnical System of Autonomous Vehicles: A Novel Application of Multi-Criteria Decision Analysis

Veljko Dubljević, George F. List, Jovan Milojevich, Nirav Ajmeri, William Bauer, Munindar P. Singh, Eleni Bardaka, Thomas Birkland, Charles Edwards, Roger Mayer, Ioan Muntean, Thomas Powers, Hesham Rakha, Vance Ricks, M. Shoaib Samandar

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.05704 2021-09-15 cs.CL cs.AI 62%

Mitigating Language-Dependent Ethnic Bias in BERT

Jaimeen Ahn, Alice Oh

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 17 pages including references and appendix. To appear in EMNLP 2021 (camera-ready ver.)

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.02032 2021-08-24 cs.CY cs.AI 62%

Socially Responsible AI Algorithms: Issues, Purposes, and Challenges

Lu Cheng, Kush R. Varshney, Huan Liu

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 45 pages, 8 figures

Journal ref Journal of Artificial Intelligence Research 71 (2021) 1137-1181

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06216 2021-08-16 cs.IR cs.AI cs.CL cs.SI 62%

MAIR: Framework for mining relationships between research articles, strategies, and regulations in the field of explainable artificial intelligence

Stanisław Gizinski, Michał Kuzba, Bartosz Pielinski, Julian Sienkiewicz, Stanisław Łaniewski, Przemysław Biecek

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.01764 2021-08-05 cs.CL cs.AI 62%

Q-Pain: A Question Answering Dataset to Measure Social Bias in Pain Management

Cécile Logé, Emily Ross, David Yaw Amoah Dadey, Saahil Jain, Adriel Saporta, Andrew Y. Ng, Pranav Rajpurkar

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to the 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.10939 2021-07-26 cs.IR cs.CY cs.LG 62%

What are you optimizing for? Aligning Recommender Systems with Human Values

Jonathan Stray, Ivan Vendrov, Jeremy Nixon, Steven Adler, Dylan Hadfield-Menell

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG

Comments Originally presented at the ICML 2020 Participatory Approaches to Machine Learning workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04243 2021-05-12 cs.CV cs.AI cs.CY 62%

Estimating and Improving Fairness with Adversarial Learning

Xiaoxiao Li, Ziteng Cui, Yifan Wu, Lin Gu, Tatsuya Harada

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 12 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.09469 2021-04-20 cs.LG cs.AI cs.HC 62%

Training Value-Aligned Reinforcement Learning Agents Using a Normative Prior

Md Sultan Al Nahian, Spencer Frazier, Brent Harrison, Mark Riedl

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments (Nahian and Frazier contributed equally to this work)

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.01685 2021-03-17 cs.AI cs.LG 62%

Agent Incentives: A Causal Perspective

Tom Everitt, Ryan Carey, Eric Langlois, Pedro A Ortega, Shane Legg

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments In Proceedings of the AAAI 2021 Conference. Supersedes arXiv:1902.09980, arXiv:2001.07118

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.10766 2021-03-01 cs.LG cs.AI stat.ML 62%

Teaching the Old Dog New Tricks: Supervised Learning with Constraints

Fabrizio Detassis, Michele Lombardi, Michela Milano

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06058 2020-12-14 cs.CY cs.AI 62%

Next Wave Artificial Intelligence: Robust, Explainable, Adaptable, Ethical, and Accountable

Odest Chadwicke Jenkins, Daniel Lopresti, Melanie Mitchell

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments A Computing Community Consortium (CCC) white paper, 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.00501 2020-03-25 cs.CY cs.AI 62%

The role of artificial intelligence in achieving the Sustainable Development Goals

Ricardo Vinuesa, Hossein Azizpour, Iolanda Leite, Madeline Balaam, Virginia Dignum, Sami Domisch, Anna Felländer, Simone Langhans, Max Tegmark, Francesco Fuso Nerini

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.11990 2020-03-16 cs.LG cs.AI stat.ML 62%

Deontological Ethics By Monotonicity Shape Constraints

Serena Wang, Maya Gupta

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments AISTATS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.12393 2020-01-17 cs.CY cs.AI physics.soc-ph 62%

To regulate or not: a social dynamics analysis of the race for AI supremacy

The Anh Han, Luis Moniz Pereira, Francisco C. Santos, Tom Lenaerts

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.11452 2019-10-28 cs.LG cs.CY stat.ML 62%

Fairness Sample Complexity and the Case for Human Intervention

Ananth Balashankar, Alyssa Lees

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG

Comments Where is the Human? Bridging the Gap Between AI and HCI, CHI Workshop 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.02134 2019-08-07 cs.CY cs.LG cs.SE 62%

Adapting SQuaRE for Quality Assessment of Artificial Intelligence Systems

Hiroshi Kuwajima, Fuyuki Ishikawa

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.06289 2019-05-16 cs.HC cs.CY cs.LG 62%

A Human-Centered Approach to Interactive Machine Learning

Kory W. Mathewson

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

Comments 4 pages, 4th Multidisciplinary Conference on Reinforcement Learning and Decision Making

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.07261 2019-02-08 cs.CY cs.AI 62%

FactSheets: Increasing Trust in AI Services through Supplier's Declarations of Conformity

Matthew Arnold, Rachel K. E. Bellamy, Michael Hind, Stephanie Houde, Sameep Mehta, Aleksandra Mojsilovic, Ravi Nair, Karthikeyan Natesan Ramamurthy, Darrell Reimer, Alexandra Olteanu, David Piorkowski, Jason Tsay, Kush R. Varshney

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.05239 2019-01-09 cs.HC cs.CY cs.LG cs.SE 62%

Improving fairness in machine learning systems: What do industry practitioners need?

Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé, Miro Dudík, Hanna Wallach

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG

Comments To appear in the 2019 ACM CHI Conference on Human Factors in Computing Systems (CHI 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.06126 2018-09-25 cs.AI cs.LG stat.ML 62%

Bootstrapping with Models: Confidence Intervals for Off-Policy Evaluation

Josiah P. Hanna, Peter Stone, Scott Niekum

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments Published in proceedings of the 16th International Conference on Autonomous Agents and Multi-agent Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.09093 2017-10-03 cs.CY cs.AI cs.HC 62%

Beyond opening up the black box: Investigating the role of algorithmic systems in Wikipedian organizational culture

R. Stuart Geiger

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 14 pages, typo fixed in v2

Journal ref Big Data & Society 4(2). 2017

详情

展开后加载摘要…

URL PDF HTML 收藏