arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1852 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1852 篇

2310.09217 2023-10-16 cs.AI 57%

Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI

Jason Hausenloy, Andrea Miotti, Claire Dennis

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05876 2023-10-10 cs.AI 57%

AI Systems of Concern

Kayla Matteucci, Shahar Avin, Fazl Barez, Seán Ó hÉigeartaigh

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 9 pages, 1 figure, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11691 2023-09-22 cs.AI cs.CR 57%

RAI4IoE: Responsible AI for Enabling the Internet of Energy

Minhui Xue, Surya Nepal, Ling Liu, Subbu Sethuvenkatraman, Xingliang Yuan, Carsten Rudolph, Ruoxi Sun, Greg Eisenhauer

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Accepted to IEEE International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15514 2023-09-12 cs.AI 57%

International Governance of Civilian AI: A Jurisdictional Certification Approach

Robert Trager, Ben Harack, Anka Reuel, Allison Carnegie, Lennart Heim, Lewis Ho, Sarah Kreps, Ranjit Lall, Owen Larter, Seán Ó hÉigeartaigh, Simon Staffell, José Jaime Villalobos

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06326 2023-09-07 cs.CY 57%

Bad, mad, and cooked: Moral responsibility for civilian harms in human-AI military teams

Susannah Kate Devitt

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 30 pages, accepted for publication in Jan Maarten Schraagen (Ed.) 'Responsible Use of AI in Military Systems', CRC Press [Forthcoming]

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02996 2023-08-08 cs.NI cs.CY cs.DC 57%

A Review of Gaps between Web 4.0 and Web 3.0 Intelligent Network Infrastructure

Zihan Zhou, Zihao Li, Xiaoshuai Zhang, Yunqing Sun, Hao Xu

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Comments 6 pages 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02025 2023-08-07 cs.CY cs.IR 57%

Applications and Societal Implications of Artificial Intelligence in Manufacturing: A Systematic Review

John P. Nelson, Justin B. Biddle, Philip Shapira

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14672 2023-06-27 stat.ML cs.LG 57%

PWSHAP: A Path-Wise Explanation Model for Targeted Variables

Lucile Ter-Minassian, Oscar Clivio, Karla Diaz-Ordaz, Robin J. Evans, Chris Holmes

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Journal ref International Conference on Machine Learning 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08959 2023-06-16 cs.CY 57%

Statutory Professions in AI governance and their consequences for explainable AI

Labhaoise NiFhaolain, Andrew Hines, Vivek Nallur

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Accepted for publication at xAI-2023 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10743 2023-06-13 cs.CY cs.HC 57%

The Ethics of AI-Generated Maps: A Study of DALLE 2 and Implications for Cartography

Yuhao Kang, Qianheng Zhang, Robert Roth

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Comments 9 pages, 3 figures, GIScience 2023 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.00796 2023-06-01 cs.LG cs.IT math.IT 57%

On Balancing Bias and Variance in Unsupervised Multi-Source-Free Domain Adaptation

Maohao Shen, Yuheng Bu, Gregory Wornell

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments ICML 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11837 2023-05-26 cs.SE cs.AI 57%

Comparing Software Developers with ChatGPT: An Empirical Investigation

Nathalia Nascimento, Paulo Alencar, Donald Cowan

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14865 2023-05-25 cs.GT cs.CY 57%

A Game-Theoretic Framework for AI Governance

Na Zhang, Kun Yue, Chao Fang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03870 2023-05-11 cs.LG 57%

Knowledge Transfer from Teachers to Learners in Growing-Batch Reinforcement Learning

Patrick Emedom-Nnamdi, Abram L. Friesen, Bobak Shahriari, Nando de Freitas, Matt W. Hoffman

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments Reincarnating Reinforcement Learning Workshop at ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04547 2023-05-09 cs.CL 57%

Diffusion Theory as a Scalpel: Detecting and Purifying Poisonous Dimensions in Pre-trained Language Models Caused by Backdoor or Bias

Zhiyuan Zhang, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

Comments Accepted by Findings of ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08001 2023-02-17 cs.AI cs.MA 57%

Learning Density-Based Correlated Equilibria for Markov Games

Libo Zhang, Yang Chen, Toru Takisaka, Bakh Khoussainov, Michael Witbrock, Jiamou Liu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12566 2023-01-31 cs.CL cs.IR 57%

Improving Cross-lingual Information Retrieval on Low-Resource Languages via Optimal Transport Distillation

Zhiqi Huang, Puxuan Yu, James Allan

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11737 2023-01-30 cs.LG 57%

Modeling human road crossing decisions as reward maximization with visual perception limitations

Yueyang Wang, Aravinda Ramakrishnan Srinivasan, Jussi P. P. Jokinen, Antti Oulasvirta, Gustav Markkula

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 6 pages, 5 figures,1 table, manuscript created for consideration at IEEE IV 2023 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.08094 2023-01-05 cs.CV cs.LG 57%

LIMEcraft: Handcrafted superpixel selection and inspection for Visual eXplanations

Weronika Hryniewska, Adrianna Grudzień, Przemysław Biecek

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Journal ref Machine Learning (2022) 1-18

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.06833 2022-11-30 cs.CY 57%

Beyond Ads: Sequential Decision-Making Algorithms in Law and Public Policy

Peter Henderson, Ben Chugg, Brandon Anderson, Daniel E. Ho

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Version 1 presented at Causal Inference Challenges in Sequential Decision Making: Bridging Theory and Practice (2021), a NeurIPS 2021 Workshop; Version 2 presented at the 2nd ACM Symposium on Computer Science and Law (2022) (DOI: https://dl.acm.org/doi/10.1145/3511265.3550439)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04491 2022-10-11 cs.LG stat.ML 57%

A survey of Identification and mitigation of Machine Learning algorithmic biases in Image Analysis

Laurent Risser, Agustin Picard, Lucas Hervier, Jean-Michel Loubes

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.10087 2022-08-23 cs.CY 57%

A Trust Framework for Government Use of Artificial Intelligence and Automated Decision Making

Pia Andrews, Tim de Sousa, Bruce Haefele, Matt Beard, Marcus Wigan, Abhinav Palia, Kathy Reid, Saket Narayan, Morgan Dumitru, Alex Morrison, Geoff Mason, Aurelie Jacquet

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Comments Comments were integrated into the paper from all peer reviewers. Am happy to provide a copied history of comments if useful

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02725 2022-08-23 cs.IR cs.CL 57%

Match-Prompt: Improving Multi-task Generalization Ability for Neural Text Matching via Prompt Learning

Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted by CIKM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08790 2022-08-19 cs.AI 57%

Explainable Reinforcement Learning on Financial Stock Trading using SHAP

Satyam Kumar, Mendhikar Vishal, Vadlamani Ravi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 28 pages; 3 Tables; 21 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.07120 2022-07-27 cs.LG cs.CV 57%

Making Corgis Important for Honeycomb Classification: Adversarial Attacks on Concept-based Explainability Tools

Davis Brown, Henry Kvinge

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments AdvML Frontiers 2022 @ ICML 2022 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.00474 2022-06-02 cs.AI cs.HC 57%

Towards Responsible AI: A Design Space Exploration of Human-Centered Artificial Intelligence User Interfaces to Investigate Fairness

Yuri Nakao, Lorenzo Strappelli, Simone Stumpf, Aisha Naseer, Daniele Regoli, Giulia Del Gamba

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 44 pages, 17 figures, the draft of a paper on International Journal of Human-Computer Interaction

Journal ref International Journal of Human-Computer Interaction, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03824 2022-05-10 cs.AI 57%

A Survey on AI Sustainability: Emerging Trends on Learning Algorithms and Research Challenges

Zhenghua Chen, Min Wu, Alvin Chan, Xiaoli Li, Yew-Soon Ong

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00656 2022-05-03 cs.CL 57%

Debiased Contrastive Learning of Unsupervised Sentence Representations

Kun Zhou, Beichen Zhang, Wayne Xin Zhao, Ji-Rong Wen

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments 11 pages, accepted by ACL 2022 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09622 2022-04-21 cs.HC cs.GL cs.LG 57%

A Brief Guide to Designing and Evaluating Human-Centered Interactive Machine Learning

Kory W. Mathewson, Patrick M. Pilarski

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 7 pages, 1 figure, Published at ML Evaluation Standards Workshop at ICLR 2022. arXiv admin note: substantial text overlap with arXiv:1905.06289

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.06933 2022-04-05 cs.LG stat.ML 57%

The reinforcement learning-based multi-agent cooperative approach for the adaptive speed regulation on a metallurgical pickling line

Anna Bogomolova, Kseniia Kingsep, Boris Voskresenskii

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏