arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1852 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1852 篇

2412.04683 2025-02-07 cs.AI 57%

From Principles to Practice: A Deep Dive into AI Ethics and Regulations

Nan Sun, Yuantian Miao, Hao Jiang, Ming Ding, Jun Zhang

机构 * University of New South Wales(新南威尔士大学) University of Newcastle(纽卡斯尔大学) Swinburne University of Technology(斯威本科技大学) Commonwealth Scientific and Industrial Research Organisation (CSIRO)(联邦科学与工业研究组织(CSIRO))

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Submitted to JAIR

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15985 2025-01-28 cs.CY 57%

Demographic Benchmarking: Bridging Socio-Technical Gaps in Bias Detection

Gemma Galdon Clavell, Rubén González-Sendino, Paola Vazquez

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15571 2025-01-28 cs.CL 57%

Cross-Cultural Fashion Design via Interactive Large Language Models and Diffusion Models

Spencer Ramsey, Amina Grant, Jeffrey Lee

机构 * Northern Caribbean University(北加勒比大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20130 2025-01-28 cs.HC cs.CY 57%

The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI Relationships

Renwen Zhang, Han Li, Han Meng, Jinyuan Zhan, Hongyuan Gan, Yi-Chieh Lee

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01957 2025-01-24 cs.AI 57%

Usage Governance Advisor: From Intent to AI Governance

Elizabeth M. Daly, Sean Rooney, Seshu Tirupathi, Luis Garces-Erice, Inge Vejsbjerg, Frank Bagehorn, Dhaval Salwala, Christopher Giblin, Mira L. Wolf-Bauwens, Ioana Giurgiu, Michael Hind, Peter Urbanetz

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 9 pages, 8 figures, AAAI workshop submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12521 2025-01-23 cs.SE cs.AI 57%

An Empirically-grounded tool for Automatic Prompt Linting and Repair: A Case Study on Bias, Vulnerability, and Optimization in Developer Prompts

Dhia Elhaq Rzig, Dhruba Jyoti Paul, Kaiser Pister, Jordan Henkel, Foyzul Hassan

机构 * University of Michigan-Dearborn(密歇根大学迪尔伯恩分校) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Microsoft(微软公司)

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.06687 2025-01-22 cs.CL 57%

Hire Me or Not? Examining Language Model's Behavior with Occupation Attributes

Damin Zhang, Yi Zhang, Geetanjali Bihani, Julia Rayz

机构 * Purdue University(普渡大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments COLING 2025

Journal ref Proceedings of the 31st International Conference on Computational Linguistics (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14062 2025-01-20 cs.HC cs.CY 57%

Understanding and Evaluating Trust in Generative AI and Large Language Models for Spreadsheets

Simon Thorne

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Journal ref Proceedings of the EuSpRIG 2024 Conference "Spreadsheet Productivity & Risks" ISBN : 978-1-905404-59-9

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06695 2025-01-14 cs.AI 57%

DVM: Towards Controllable LLM Agents in Social Deduction Games

Zheng Zhang, Yihuai Lan, Yangsen Chen, Lei Wang, Xiang Wang, Hao Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Singapore Management University(新加坡管理大学) University of Science and Technology of China(中国科学技术大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17114 2025-01-14 cs.AI cs.ET 57%

Decentralized Governance of Autonomous AI Agents

Tomer Jordi Chaffer, Charles von Goins, Bayo Okusanya, Dontrail Cotlage, Justin Goldston

机构 * Gemach DAO Rochester Institute of Technology(罗切斯特理工学院) NPC Labs(NPC实验室) National University(国立大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05617 2025-01-13 cs.CY cs.DL 57%

Datasheets for Healthcare AI: A Framework for Transparency and Bias Mitigation

Marjia Siddik, Harshvardhan J. Pandit

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments Irish Conference on Artificial Intelligence and Cognitive Science (AICS), December 2024, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04437 2025-01-09 eess.SY cs.AI cs.ET cs.SY 57%

Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions

Doaa Mahmud, Hadeel Hajmohamed, Shamma Almentheri, Shamma Alqaydi, Lameya Aldhaheri, Ruhul Amin Khalil, Nasir Saeed

机构 * College of Engineering, UAE University(阿联酋大学工程学院) Engineering Requirement Unit (ERU), College of Engineering, UAE University(阿联酋大学工程学院工程需求单元(ERU))

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Accepted for publication in IEEE Transactions on Intelligent Transportation Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19915 2025-01-08 econ.GN cs.AI q-fin.EC 57%

AI-Driven Scenarios for Urban Mobility: Quantifying the Role of ODE Models and Scenario Planning in Reducing Traffic Congestion

Katsiaryna Bahamazava

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02368 2025-01-07 cs.AI cs.HC 57%

Enhancing Workplace Productivity and Well-being Using AI Agent

Ravirajan K, Arvind Sundarajan

机构 * LTIMindtree

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18588 2024-12-25 cs.RO cs.AI cs.SY eess.SY 57%

A Paragraph is All It Takes: Rich Robot Behaviors from Interacting, Trusted LLMs

OpenMind, Shaohong Zhong, Adam Zhou, Boyuan Chen, Homin Luo, Jan Liphardt

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17505 2024-12-24 stat.ML cs.LG 57%

More is Less? A Simulation-Based Approach to Dynamic Interactions between Biases in Multimodal Models

Mounia Drissi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00469 2024-12-23 cs.CY 57%

Beyond Incompatibility: Trade-offs between Mutually Exclusive Fairness Criteria in Machine Learning and Law

Meike Zehlike, Alex Loosley, Håkan Jonsson, Emil Wiedemann, Philipp Hacker

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04310 2024-12-19 cs.CL 57%

Montague semantics and modifier consistency measurement in neural language models

Danilo S. Carvalho, Edoardo Manino, Julia Rozanova, Lucas Cordeiro, André Freitas

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11335 2024-12-17 cs.CY 57%

Generative AI regulation can learn from social media regulation

Ruth Elisabeth Appel

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Presented at 2nd Workshop on Regulatable ML at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09799 2024-12-16 cs.CV cs.AI 57%

CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection

Qibo Chen, Weizhong Jin, Jianyue Ge, Mengdi Liu, Yuchao Yan, Jian Jiang, Li Yu, Xuanjiang Guo, Shuchang Li, Jianzhong Chen

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07880 2024-12-13 cs.AI 57%

Towards Foundation-model-based Multiagent System to Accelerate AI for Social Impact

Yunfan Zhao, Niclas Boehmer, Aparna Taneja, Milind Tambe

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08181 2024-12-05 cs.AI 57%

Challenges in Guardrailing Large Language Models for Science

Nishan Pantha, Muthukumaran Ramasubramanian, Iksha Gurung, Manil Maskey, Rahul Ramachandran

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16691 2024-11-27 cs.CY 57%

The Dual Impact of Artificial Intelligence in Healthcare: Balancing Advancements with Ethical and Operational Challenges

Balaji Shesharao Ingole, Vishnu Ramineni, Nikhil Kumar Pulipeta, Manoj Jayntilal Kathiriya, Manjunatha Sughaturu Krishnappa, Vivekananda Jayaram

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Journal ref European Journal of Computer Science and Information Technology,12 (6),35-45, 2024 Print ISSN: 2054-0957 (Print)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13659 2024-11-27 cs.AI 57%

Leveraging Large Language Models for Patient Engagement: The Power of Conversational AI in Digital Health

Bo Wen, Raquel Norel, Julia Liu, Thaddeus Stappenbeck, Farhana Zulkernine, Huamin Chen

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 10 pages, 6 figures, ICDH 2024 invited paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16193 2024-11-26 cs.CY 57%

The Critical Canvas--How to regain information autonomy in the AI era

Dong Chen

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00434 2024-11-25 cs.CY cs.RO 57%

Rapid Integration of LLMs in Healthcare Raises Ethical Concerns: An Investigation into Deceptive Patterns in Social Robots

Robert Ranisch, Joschka Haltaufderheide

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 7 pages, 1table, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08896 2024-11-15 eess.SP cs.LG cs.NI 57%

Demand-Aware Beam Hopping and Power Allocation for Load Balancing in Digital Twin empowered LEO Satellite Networks

Ruili Zhao, Jun Cai, Jiangtao Luo, Junpeng Gao, Yongyi Ran

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02577 2024-11-06 cs.CY 57%

Where Assessment Validation and Responsible AI Meet

Jill Burstein, Geoffrey T. LaFlair

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00860 2024-11-05 cs.CL cs.CV 57%

Survey of Cultural Awareness in Language Models: Text and Beyond

Siddhesh Pawar, Junyeong Park, Jiho Jin, Arnav Arora, Junho Myung, Srishti Yadav, Faiz Ghifari Haznitrama, Inhwa Song, Alice Oh, Isabelle Augenstein

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23310 2024-11-01 q-bio.NC cs.AI 57%

Moral Agency in Silico: Exploring Free Will in Large Language Models

Morgan S. Porter

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏