arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2504.00652 2025-04-02 cs.CY cs.AI cs.ET 62%

Towards Adaptive AI Governance: Comparative Insights from the U.S., EU, and Asia

Vikram Kulothungan, Deepti Gupta

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Accepted at IEEE BigDataSecurity 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00241 2025-04-02 cs.CL cs.AI 62%

Synthesizing Public Opinions with LLMs: Role Creation, Impacts, and the Future to eDemorcacy

Rabimba Karanjai, Boris Shor, Amanda Austin, Ryan Kennedy, Yang Lu, Lei Xu, Weidong Shi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24228 2025-04-01 cs.AI cs.CL cs.MA 62%

PAARS: Persona Aligned Agentic Retail Shoppers

Saab Mansour, Leonardo Perelli, Lorenzo Mainetti, George Davidson, Stefano D'Amato

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17756 2025-03-25 cs.LG cs.AI cs.CR cs.NE 62%

Bandwidth Reservation for Time-Critical Vehicular Applications: A Multi-Operator Environment

Abdullah Al-Khatib, Abdullah Ahmed, Klaus Moessner, Holger Timinger

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06795 2025-03-13 cs.CL cs.AI 62%

Bridging the Fairness Gap: Enhancing Pre-trained Models with LLM-Generated Sentences

Liu Yu, Ludie Guo, Ping Kuang, Fan Zhou

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14210 2025-03-13 cs.LG cs.AI 62%

Fair Overlap Number of Balls (Fair-ONB): A Data-Morphology-based Undersampling Method for Bias Reduction

José Daniel Pascual-Triana, Alberto Fernández, Paulo Novais, Francisco Herrera

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 14 pages, 5 tables, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06523 2025-03-11 cs.CY cs.AI 62%

Generative AI as Digital Media

Gilad Abiri

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Journal ref Harv. J. Sports & Ent. L. 15 (2024): 279

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05937 2025-03-11 cs.CY cs.AI 62%

The Unified Control Framework: Establishing a Common Foundation for Enterprise AI Governance, Risk Management and Regulatory Compliance

Ian W. Eisenberg, Lucía Gamboa, Eli Sherman

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02255 2025-03-05 cs.CL cs.LG 62%

AxBERT: An Interpretable Chinese Spelling Correction Method Driven by Associative Knowledge Network

Fanyu Wang, Hangyu Zhu, Zhenping Xie

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16696 2025-02-25 cs.LG cs.AI 62%

Dynamic LLM Routing and Selection based on User Preferences: Balancing Performance, Cost, and Ethics

Deepak Babu Piskala, Vijay Raajaa, Sachin Mishra, Bruno Bozza

专题命中 AI治理与伦理 :harmlessness(abstract);分类 cs.AI、cs.LG

Journal ref International Journal of Computer Applications, Vol. 186, No. 51, November 2024, pp. 1-7

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15357 2025-02-24 cs.CY cs.AI 62%

Integrating Generative AI in Cybersecurity Education: Case Study Insights on Pedagogical Strategies, Critical Thinking, and Responsible AI Use

Mahmoud Elkhodr, Ergun Gide

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03429 2025-02-06 cs.CL cs.AI 62%

On Fairness of Unified Multimodal Large Language Model for Image Generation

Ming Liu, Hao Chen, Jindong Wang, Liwen Wang, Bhiksha Raj Ramakrishnan, Wensheng Zhang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00011 2025-02-04 cs.CY cs.AI cs.HC 62%

TOAST Framework: A Multidimensional Approach to Ethical and Sustainable AI Integration in Organizations

Dian Tjondronegoro

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 25 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15287 2025-01-30 cs.CL cs.AI 62%

Sycophancy in Large Language Models: Causes and Mitigations

Lars Malmqvist

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref Computing Conference 2025 (upcoming)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04472 2025-01-30 cs.CL cs.CY 62%

Collapsed Language Models Promote Fairness

Jingxuan Xu, Wuyang Chen, Linyi Li, Yao Zhao, Yunchao Wei

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17420 2025-01-30 cs.CL cs.AI cs.HC 62%

Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models

Yuxuan Li, Hirokazu Shirado, Sauvik Das

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11820 2025-01-23 cs.CY cs.AI 62%

Responsible AI Question Bank: A Comprehensive Tool for AI Risk Assessment

Sung Une Lee, Harsha Perera, Yue Liu, Boming Xia, Qinghua Lu, Liming Zhu, Olivier Salvado, Jon Whittle

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 30 pages, 6 tables, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10391 2025-01-22 cs.CY cs.AI 62%

Developing an Ontology for AI Act Fundamental Rights Impact Assessments

Tytti Rintamaki, Harshvardhan J. Pandit

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments Presented at CLAIRvoyant (ConventicLE on Artificial Intelligence Regulation) Workshop 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10390 2025-01-22 cs.CY cs.AI 62%

Towards an Environmental Ethics of Artificial Intelligence

Nynke van Uffelen, Lode Lauwaert, Mark Coeckelbergh, Olya Kudina

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09182 2025-01-17 cs.AI cs.CR cs.CY cs.SE 62%

A Blockchain-Enabled Approach to Cross-Border Compliance and Trust

Vikram Kulothungan

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments This is a preprint of paper that has been accepted for Publication at 2024 IEEE International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01973 2025-01-10 cs.CV cs.AI cs.CY 62%

INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models

Di Jin, Xing Liu, Yu Liu, Jia Qing Yap, Andrea Wong, Adriana Crespo, Qi Lin, Zhiyuan Yin, Qiang Yan, Ryan Ye

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Di Jin and Xing Liu contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14571 2025-01-10 cs.HC cs.AI cs.CY 62%

Driving Towards Inclusion: A Systematic Review of AI-powered Accessibility Enhancements for People with Disability in Autonomous Vehicles

Ashish Bastola, Hao Wang, Sayed Pedram Haeri Boroujeni, Julian Brinkley, Ata Jahangir Moshayedi, Abolfazl Razi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05937 2024-12-10 cs.LG cs.AI cs.IR cs.MA 62%

Accelerating Manufacturing Scale-Up from Material Discovery Using Agentic Web Navigation and Retrieval-Augmented AI for Process Engineering Schematics Design

Sakhinana Sagar Srinivas, Akash Das, Shivam Gupta, Venkataramana Runkana

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19308 2024-12-10 cs.CL cs.AI 62%

Designing Domain-Specific Large Language Models: The Critical Role of Fine-Tuning in Public Opinion Simulation

Haocheng Lin

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01075 2024-12-03 cs.LG cs.AI 62%

Multi-Agent Deep Reinforcement Learning for Distributed and Autonomous Platoon Coordination via Speed-regulation over Large-scale Transportation Networks

Dixiao Wei, Peng Yi, Jinlong Lei, Xingyi Zhu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00962 2024-12-03 cs.AI cs.CL cs.SC 62%

LLMs as mirrors of societal moral standards: reflection of cultural divergence and agreement across ethical topics

Mijntje Meijer, Hadi Mohammadi, Ayoub Bagheri

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02885 2024-12-03 cs.HC cs.CL cs.CY cs.SI 62%

CogErgLLM: Exploring Large Language Model Systems Design Perspective Using Cognitive Ergonomics

Azmine Toushik Wasi, Mst Rafia Islam

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.CY

Comments 10 Page, 3 Figures. Accepted in: (i) ICML'24: LLMs & Cognition Workshop (Non-archival; OpenReview: https://openreview.net/forum?id=63C9YSc77p) (ii) EMNLP'24 : NLP for Science Workshop (Archival; ACL Anthology: https://aclanthology.org/2024.nlp4science-1.22/)

Journal ref Proceedings of the 1st Workshop on NLP for Science (NLP4Science), EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20739 2024-12-02 cs.CL cs.AI 62%

Gender Bias in LLM-generated Interview Responses

Haein Kong, Yongsu Ahn, Sangyub Lee, Yunho Maeng

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to NeurlIPS 2024, SoLaR workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.05283 2024-12-02 cs.CL cs.AI 62%

On the Relationship between Truth and Political Bias in Language Models

Suyash Fulay, William Brannon, Shrestha Mohanty, Cassandra Overney, Elinor Poole-Dayan, Deb Roy, Jad Kabbara

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024

Journal ref Proc. EMNLP (2024) 9004-9018

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12762 2024-11-27 cs.AI cs.CY 62%

How should AI decisions be explained? Requirements for Explanations from the Perspective of European Law

Benjamin Fresz, Elena Dubovitskaya, Danilo Brajovic, Marco Huber, Christian Horz

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Journal ref Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (2024), 7(1), 438-450

详情

展开后加载摘要…

URL PDF HTML 收藏