arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1852 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1852 篇

2401.03093 2024-05-21 cs.AI q-bio.PE 57%

XXAI: Towards eXplicitly eXplainable Artificial Intelligence

V. L. Kalmykov, L. V. Kalmykov

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 18 pages, 1 graphical abstract, 1 figure, 72 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13765 2024-05-21 cs.AI 57%

Towards ethical multimodal systems

Alexis Roger, Esma Aïmeur, Irina Rish

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 5 pages, multimodal ethical dataset building, accepted in the NeurIPS 2023 MP2 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13605 2024-05-07 cs.CY 57%

Regulating AI-Based Remote Biometric Identification. Investigating the Public Demand for Bans, Audits, and Public Database Registrations

Kimon Kieslich, Marco Lünich

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02846 2024-05-07 cs.AI 57%

Responsible AI: Portraits with Intelligent Bibliometrics

Yi Zhang, Mengjia Wu, Guangquan Zhang, Jie Lu

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01697 2024-05-06 cs.CY 57%

Towards an Ethical and Inclusive Implementation of Artificial Intelligence in Organizations: A Multidimensional Framework

Ernesto Giralt Hernández

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments This is an English version of the original article arXiv:2405.00225v1 [cs.CY] (Hacia una implementación ética e inclusiva de la Inteligencia Artificial en las organizaciones: un marco multidimensional)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00225 2024-05-02 cs.CY 57%

Hacia una implementación ética e inclusiva de la Inteligencia Artificial en las organizaciones: un marco multidimensional

Ernesto Giralt Hernández

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments in Spanish language

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17522 2024-04-29 cs.SE cs.AI 57%

Enhancing Legal Compliance and Regulation Analysis with Large Language Models

Shabnam Hassani

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments to be published in 32nd IEEE International Requirements Engineering 2024 Conference (RE'24) - Doctoral Symposium. arXiv admin note: text overlap with arXiv:2404.14356

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05414 2024-03-26 cs.RO cs.AI 57%

Ethics of Artificial Intelligence and Robotics in the Architecture, Engineering, and Construction Industry

Ci-Jyun Liang, Thai-Hoa Le, Youngjib Ham, Bharadwaj R. K. Mantha, Marvin H. Cheng, Jacob J. Lin

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 109 pages, 5 figures, submitted to Automation in Construction

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14049 2024-03-22 cs.RO cs.AI cs.HC 57%

A Roadmap Towards Automated and Regulated Robotic Systems

Yihao Liu, Mehran Armand

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12624 2024-03-20 cs.LG math.OC stat.ML 57%

Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent Stochastic Approach

Heshan Fernando, Han Shen, Miao Liu, Subhajit Chaudhury, Keerthiram Murugesan, Tianyi Chen

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments Changed hyper-parameter choice which affects some of the convergence rate results in the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05696 2024-03-12 cs.CL cs.CV 57%

SeeGULL Multilingual: a Dataset of Geo-Culturally Situated Stereotypes

Mukul Bhutani, Kevin Robinson, Vinodkumar Prabhakaran, Shachi Dave, Sunipa Dev

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00292 2024-03-07 cs.CY 57%

Sustainable AI Regulation

Philipp Hacker

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Comments Privacy Law Scholars Conference 2023; Common Market Law Review (forthcoming)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11190 2024-02-20 cs.CL 57%

Disclosure and Mitigation of Gender Bias in LLMs

Xiangjue Dong, Yibo Wang, Philip S. Yu, James Caverlee

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments The first two authors contribute equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09328 2024-02-15 stat.ML cs.LG stat.ME 57%

Connecting Algorithmic Fairness to Quality Dimensions in Machine Learning in Official Statistics and Survey Production

Patrick Oliver Schenk, Christoph Kern

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08797 2024-02-15 cs.CY 57%

Computing Power and the Governance of Artificial Intelligence

Girish Sastry, Lennart Heim, Haydn Belfield, Markus Anderljung, Miles Brundage, Julian Hazell, Cullen O'Keefe, Gillian K. Hadfield, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Janet Egan, Robert F. Trager, Shahar Avin, Adrian Weller, Yoshua Bengio, Diane Coyle

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Figures can be accessed at: https://github.com/lheim/CPGAI-Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04017 2024-02-13 cs.LG q-bio.QM 57%

PGraphDTA: Improving Drug Target Interaction Prediction using Protein Language Models and Contact Maps

Rakesh Bal, Yijia Xiao, Wei Wang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments AI for Science Workshop, NeurIPS 2023. 11 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05786 2024-02-12 cs.AI cs.GT 57%

Prompting Fairness: Artificial Intelligence as Game Players

Jazmia Henry

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00745 2024-02-02 cs.CL 57%

Enhancing Ethical Explanations of Large Language Models through Iterative Symbolic Refinement

Xin Quan, Marco Valentino, Louise A. Dennis, André Freitas

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Camera-ready for EACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15798 2024-01-30 cs.CL 57%

UnMASKed: Quantifying Gender Biases in Masked Language Models through Linguistically Informed Job Market Prompts

Iñigo Parra

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments EACL 2024 SRW

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09852 2024-01-19 cs.CV cs.AI 57%

Enhancing the Fairness and Performance of Edge Cameras with Explainable AI

Truong Thanh Hung Nguyen, Vo Thanh Khang Nguyen, Quoc Hung Cao, Van Binh Truong, Quoc Khanh Nguyen, Hung Cao

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments IEEE ICCE 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18007 2023-12-01 astro-ph.IM astro-ph.GA cs.LG 57%

Towards out-of-distribution generalization in large-scale astronomical surveys: robust networks learn similar representations

Yash Gondhalekar, Sultan Hassan, Naomi Saphra, Sambatra Andrianomena

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments Accepted to Machine Learning and the Physical Sciences Workshop, NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04955 2023-11-17 cs.LG 57%

Information-Theoretic Bounds on The Removal of Attribute-Specific Bias From Neural Networks

Jiazhi Li, Mahyar Khayatkhoei, Jiageng Zhu, Hanchen Xie, Mohamed E. Hussein, Wael AbdAlmageed

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

Comments 15 pages, 4 figures, 3 tables. To appear in Algorithmic Fairness through the Lens of Time Workshop at NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09687 2023-11-17 cs.CL 57%

Inducing Political Bias Allows Language Models Anticipate Partisan Reactions to Controversies

Zihao He, Siyi Guo, Ashwin Rao, Kristina Lerman

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16607 2023-11-14 cs.CL 57%

On the Interplay between Fairness and Explainability

Stephanie Brandl, Emanuele Bugliarello, Ilias Chalkidis

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL

Comments 15 pages (incl Appendix), 4 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18130 2023-11-09 cs.CL cs.HC 57%

DELPHI: Data for Evaluating LLMs' Performance in Handling Controversial Issues

David Q. Sun, Artem Abzaliev, Hadas Kotek, Zidi Xiu, Christopher Klein, Jason D. Williams

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

Comments Accepted to EMNLP Industry Track 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04411 2023-11-08 cs.LG 57%

Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL

Yang Yue, Rui Lu, Bingyi Kang, Shiji Song, Gao Huang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments 31 pages, 20 figures

Journal ref NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01550 2023-11-06 cs.AI econ.GN q-fin.EC 57%

Market Concentration Implications of Foundation Models

Jai Vipra, Anton Korinek

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Working Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11090 2023-11-01 cs.CV cs.LG stat.AP 57%

Fairness Explainability using Optimal Transport with Applications in Image Classification

Philipp Ratz, François Hu, Arthur Charpentier

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18832 2023-10-31 cs.AI 57%

Responsible AI (RAI) Games and Ensembles

Yash Gupta, Runtian Zhai, Arun Suggala, Pradeep Ravikumar

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16360 2023-10-26 cs.AI cs.RO 57%

A Comprehensive Review of AI-enabled Unmanned Aerial Vehicle: Trends, Vision , and Challenges

Osim Kumar Pal, Md Sakib Hossain Shovon, M. F. Mridha, Jungpil Shin

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏