arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1852 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1852 篇

2511.14693 2025-11-19 cs.CL 57%

Talk, Snap, Complain: Validation-Aware Multimodal Expert Framework for Fine-Grained Customer Grievances

Rishu Kumar Singh, Navneet Shreya, Sarmistha Das, Apoorva Singh, Sriparna Saha

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments To be published in the Proceedings of the 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026 Special Track on AI for Social Impact )

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13865 2025-11-19 econ.GN cs.AI q-fin.EC 57%

Randomized Controlled Trials for Conditional Access Optimization Agent

James Bono, Beibei Cheng, Joaquin Lozano

机构 * Microsoft Corporation(微软公司)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11789 2025-11-18 cs.MA cs.AI 57%

From Single to Societal: Analyzing Persona-Induced Bias in Multi-Agent Interactions

Jiayi Li, Xiao Liu, Yansong Feng

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments AAAI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08641 2025-11-13 cs.CR cs.CY cs.MA 57%

QOC DAO -- Stepwise Development Towards an AI Driven Decentralized Autonomous Organization

Marc Jansen, Christophe Verdot

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07941 2025-11-12 cs.CV cs.AI 57%

Libra-MIL: Multimodal Prototypes Stereoscopic Infused with Task-specific Language Priors for Few-shot Whole Slide Image Classification

Zhenfeng Zhuang, Fangyu Zhou, Liansheng Wang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02241 2025-11-12 cs.CL 57%

Isolating Culture Neurons in Multilingual Large Language Models

Danial Namazifard, Lukas Galke Poech

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted at IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23628 2025-11-11 physics.soc-ph cs.CY cs.DS cs.GT econ.TH 57%

Matchings Under Biased and Correlated Evaluations

Amit Kumar, Nisheeth K. Vishnoi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments To appear in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13182 2025-11-11 cs.LO cs.AI 57%

Information Science Principles of Machine Learning: A Causal Chain Meta-Framework Based on Formalized Information Mapping

Jianfeng Xu

机构 * Koguan School of Law, China Institute for Smart Justice, School of Computer Science, Shanghai Jiao Tong University(柯 guar 法学院、中国智能正义研究院、计算机科学学院、上海交通大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25863 2025-11-05 cs.CR cs.AI 57%

AAGATE: A NIST AI RMF-Aligned Governance Platform for Agentic AI

Ken Huang, Kyriakos Rock Lambros, Jerry Huang, Yasir Mehmood, Hammad Atta, Joshua Beck, Vineeth Sai Narajala, Muhammad Zeeshan Baig, Muhammad Aziz Ul Haq, Nadeem Shahzad, Bhavya Gupta

机构 * RockCyber Kleiner Perkins Qorvex Consulting & Roshan Consulting SAS Institute OWASP Wentworth Institute of Higher Education & Machine Learning Professional Skylink Antenna Roshan Consulting & Robotic Process Automation Stanford University

专题命中 AI治理与伦理 :red teaming(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01550 2025-11-04 cs.AI 57%

Analyzing Sustainability Messaging in Large-Scale Corporate Social Media

Ujjwal Sharma, Stevan Rudinac, Ana Mićković, Willemijn van Dolen, Marcel Worring

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26309 2025-10-31 cs.AI cs.IR 57%

GraphCompliance: Aligning Policy and Context Graphs for LLM-Based Regulatory Compliance

Jiseong Chung, Ronny Ko, Wonchul Yoo, Makoto Onizuka, Sungmok Kim, Tae-Wan Kim, Won-Yong Shin

机构 * Seoul National University(首尔国立大学) Osaka University(大阪大学) Yonsei University(延世大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments Under review at The Web Conference 2026 (Semantics & Knowledge track). Code will be released upon acceptance. This arXiv v1 contains no repository links to preserve double-blind review

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10492 2025-10-29 cs.LG 57%

Blockchain-Based Federated Learning: Incentivizing Data Sharing and Penalizing Dishonest Behavior

Amir Jaberzadeh, Ajay Kumar Shrestha, Faijan Ahamad Khan, Mohammed Afaan Shaikh, Bhargav Dave, Jason Geng

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

Comments To appear in the 5th International Congress on Blockchain and Applications (BLOCKCHAIN'23). Publish by the Lecture Notes in Networks and Systems series of Springer Verlag

Journal ref Blockchain and Applications, 5th International Congress. BLOCKCHAIN 2023. Lecture Notes in Networks and Systems, vol 778. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22286 2025-10-28 cs.CY 57%

Hybrid Instructor Ai Assessment In Academic Projects: Efficiency, Equity, And Methodological Lessons

Hugo Roger Paz

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments 12 pages, in Spanish language, 0 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21967 2025-10-28 cs.HC cs.CY 57%

We Need Accountability in Human-AI Agent Relationships

Benjamin Lange, Geoff Keeling, Arianna Manzini, Amanda McCroskery

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21203 2025-10-27 cs.CY 57%

The Nuclear Analogy in AI Governance Research

Sophia Hatz

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Hatz, S. (in press). The Nuclear Analogy in AI Governance Research. In M. Furendal & M. Lundgren (Eds.), Handbook on the Global Governance of Artificial Intelligence. Edward Elgar Publishing

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19196 2025-10-23 cs.CY 57%

Integration of AI in STEM Education, Addressing Ethical Challenges in K-12 Settings

Shaouna Shoaib Lodhi, Shoaib Lodhi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments This paper pursues three goals: (1) analyzing ethical challenges in AI-driven STEM education, (2) evaluating AI and ethics curricula for STEM relevance, and (3) proposing a research-based framework for responsible integration that bridges teacher readiness gaps and promotes equity through STEM-focused strategies

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18556 2025-10-22 cs.CL 57%

Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency

Svetlana Maslenkova, Clement Christophe, Marco AF Pimentel, Tathagata Raha, Muhammad Umar Salman, Ahmed Al Mahrooqi, Avani Gupta, Shadab Khan, Ronnie Rajan, Praveenkumar Kanithi

机构 * M42, Abu Dhabi(阿布扎比M42)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

Comments Accepted to EMNLP Main 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04650 2025-10-21 cs.LG 57%

Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction

Zesheng Ye, Chengyi Cai, Ruijiang Dong, Jianzhong Qi, Lei Feng, Pin-Yu Chen, Feng Liu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02649 2025-10-21 cs.AI 57%

Fully Autonomous AI Agents Should Not be Developed

Margaret Mitchell, Avijit Ghosh, Alexandra Sasha Luccioni, Giada Pistilli

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14443 2025-10-17 cs.SD cs.AI eess.AS 57%

Big Data Approaches to Bovine Bioacoustics: A FAIR-Compliant Dataset and Scalable ML Framework for Precision Livestock Welfare

Mayuri Kate, Suresh Neethirajan

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 40 pages, 14 figures, 9 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10623 2025-10-17 cs.LG cs.CV 57%

Flows and Diffusions on the Neural Manifold

Daniel Saragih, Deyu Cao, Tejas Balaji

机构 * Queen’s University and Vector Institute(女王大学和向量研究所) University of Tokyo(东京大学) University of Toronto(多伦多大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 43 pages, 11 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13699 2025-10-17 cs.IR cs.AI 57%

A Comprehensive Review of Recommender Systems: Transitioning from Theory to Practice

Shaina Raza, Mizanur Rahman, Safiullah Kamawal, Armin Toroghi, Ananya Raval, Farshad Navah, Amirmohammad Kazemeini

机构 * Vector Institute(向量研究所) Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments we quarterly update of this literature

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13195 2025-10-16 cs.AI 57%

Emotional Cognitive Modeling Framework with Desire-Driven Objective Optimization for LLM-empowered Agent in Social Simulation

Qun Ma, Xiao Xue, Xuwen Zhang, Zihan Zhao, Yuwei Guo, Ming Zhang

机构 * College of Intelligence and Computing(智能与计算学院) Faculty of Environment, Science and Economy(环境、科学与经济学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16152 2025-10-16 cs.CL 57%

Towards Region-aware Bias Evaluation Metrics

Angana Borah, Aparna Garimella, Rada Mihalcea

机构 * University of Michigan(密歇根大学) Adobe Research(Adobe研究)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted to Cross-Cultural Considerations in NLP (C3NLP Workshop at NAACL 2025) -- Outstanding Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18943 2025-10-15 cs.CL 57%

MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems

Xuanming Zhang, Yuxuan Chen, Samuel Yeh, Sharon Li

机构 * Tsinghua University(清华大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11595 2025-10-14 cs.AI cs.GL 57%

Reproducibility: The New Frontier in AI Governance

Israel Mason-Williams, Gabryel Mason-Williams

机构 * Imperial College London(帝国理工学院) King's College London(国王学院) Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 12 pages,6 figures,Workshop on Technical AI Governance at ICML

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11079 2025-10-14 cs.AI 57%

Argumentation-Based Explainability for Legal AI: Comparative and Regulatory Perspectives

Andrada Iulia Prajescu, Roberto Confalonieri

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08086 2025-10-10 cs.AI 57%

From Ethical Declarations to Provable Independence: An Ontology-Driven Optimal-Transport Framework for Certifiably Fair AI Systems

Sukriti Bhattacharya, Chitro Majumdar

机构 * Senior Scientist, Trustworthy AI, Luxembourg Institute of Science \& Technology, Maison de l'innovation 5, L-4362, Luxembourg Chief Investment Risk Strategist for Sovereign Institutions Founder, RsRL, Jumeirah Beach Residence (JBR), P.O.Box 29215, Dubai, United Arab Emirates

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 19 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06623 2025-10-09 cs.LG 57%

DPA-Net: A Dual-Path Attention Neural Network for Inferring Glycemic Control Metrics from Self-Monitored Blood Glucose Data

Canyu Lei, Benjamin Lobo, Jianxin Xie

机构 * Binghamton University, Department of Computer Science(宾夕法尼亚州立大学比恩代特分校计算机科学系) University of Virginia, School of Data Science(弗吉尼亚大学数据科学学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04577 2025-10-07 cs.SD cs.LG cs.MM eess.AS 57%

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

Juncheng Wang, Chao Xu, Cheng Yu, Zhe Hu, Haoyu Xie, Guoqi Yu, Lei Shang, Shujun Wang

机构 * The Hong Kong Polytechnic University(香港理工大学) Alibaba Group(阿里巴巴集团)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏