arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1852 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1852 篇

1905.10985 2020-02-04 cs.AI 57%

AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence

Jeff Clune

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.01723 2019-07-04 stat.ML cs.LG stat.AP 57%

Towards Interpretable Deep Extreme Multi-label Learning

Yihuang Kang, I-Ling Cheng, Wenjui Mao, Bowen Kuo, Pei-Ju Lee

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.04585 2019-02-26 cs.AI q-fin.GN stat.ML 57%

Categorizing Variants of Goodhart's Law

David Manheim, Scott Garrabrant

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.08221 2019-01-25 cs.AI 57%

When is it right and good for an intelligent autonomous vehicle to take over control (and hand it back)?

Ajit Narayanan

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.06064 2018-07-18 cs.LG stat.ML 57%

Online Robust Policy Learning in the Presence of Unknown Adversaries

Aaron J. Havens, Zhanhong Jiang, Soumik Sarkar

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 18 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.10322 2018-06-28 cs.AI 57%

The Virtuous Machine - Old Ethics for New Technology?

Nicolas Berberich, Klaus Diepold

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.10820 2018-05-29 cs.AI 57%

Local Rule-Based Explanations of Black Box Decision Systems

Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Dino Pedreschi, Franco Turini, Fosca Giannotti

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.09854 2018-01-31 cs.AI 57%

Algorithms for the Greater Good! On Mental Modeling and Acceptable Symbiosis in Human-AI Collaboration

Tathagata Chakraborti, Subbarao Kambhampati

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.06035 2017-11-17 cs.AI 57%

From Algorithmic Black Boxes to Adaptive White Boxes: Declarative Decision-Theoretic Ethical Programs as Codes of Ethics

Martijn van Otterlo

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 7 pages, 1 figure, submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
1504.03592 2015-04-15 cs.AI 57%

Towards Verifiably Ethical Robot Behaviour

Louise A. Dennis, Michael Fisher, Alan F. T. Winfield

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Presented at the 1st International Workshop on AI and Ethics, Sunday 25th January 2015, Hill Country A, Hyatt Regency Austin. Will appear in the workshop proceedings published by AAAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02992 2025-10-27 cs.AI cs.CL cs.LG 56%

Mitigating Manipulation and Enhancing Persuasion: A Reflective Multi-Agent Approach for Legal Argument Generation

Li Zhang, Kevin D. Ashley

机构 * Intelligent Systems Program University of Pittsburgh Pittsburgh Pennsylvania USA Intelligent Systems Program University of Pittsburgh

专题命中 AI治理与伦理 :分类 cs.CL、cs.AI、cs.LG;safety(comments)

Comments 13 pages, 2 figures, 2nd ConventicLe on Artificial Intelligence Regulation and Safety Workshop at ICAIL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02386 2024-08-06 cs.HC 56%

Responsibility and Regulation: Exploring Social Measures of Trust in Medical AI

Glenn McGarry, Andy Crabtree, Lachlan Urquhart, Alan Chamberlain

专题命中 AI治理与伦理 :trustworthy(abstract,comments)

Comments To be published in Second International Symposium on Trustworthy Autonomous Systems, September 15 18, 2024, Austin, Texas

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05307 2023-06-09 cs.CL cs.CY cs.LG stat.ML 56%

Are fairness metric scores enough to assess discrimination biases in machine learning?

Fanny Jourdan, Laurent Risser, Jean-Michel Loubes, Nicholas Asher

专题命中 AI治理与伦理 :分类 cs.CL、cs.CY、cs.LG;trustworthy(comments)

Comments Accepted for publication at Third Workshop on Trustworthy Natural Language Processing, ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14402 2022-11-29 cs.CL cs.CY cs.LG 56%

An Analysis of Social Biases Present in BERT Variants Across Multiple Languages

Aristides Milios, Parishad BehnamGhader

专题命中 AI治理与伦理 :分类 cs.CL、cs.CY、cs.LG;trustworthy(comments)

Comments Accepted to 2022 Trustworthy and Socially Responsible Machine Learning (TSRML 2022) Workshop at NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13130 2022-11-24 cs.CY cs.AI cs.LG 56%

A Brief Overview of AI Governance for Responsible Machine Learning Systems

Navdeep Gill, Abhishek Mathur, Marcos V. Conde

专题命中 AI治理与伦理 :分类 cs.AI、cs.CY、cs.LG;trustworthy(comments)

Comments NeurIPS 2022 Trustworthy and Socially Responsible Machine Learning (TSRML) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10028 2026-03-12 cs.CY cs.AI 54%

How to Count AIs: Individuation and Liability for AI Agents

如何计算AI:AI代理的个体化与责任归属

Yonathan Arbel, Peter Salib, Simon Goldstein

机构 * University of Alabama(阿拉巴马大学) University of Hong Kong(香港大学) University of Houston(休斯顿大学) HKU AI & Humanity Lab(香港大学AI与人类实验室) Center for Law & AI Risk(法律与人工智能风险中心) Institute for Law & AI(法律与人工智能研究所)

专题命中 AI治理与伦理 :分类 cs.AI、cs.CY;safety(comments);AI safety(comments)

AI总结 本文提出通过‘算法公司’解决AI代理的个体识别与责任归属问题,通过法律虚构实体实现对AI行为的追踪与责任划分。

Comments 36 pages. Presented at the Law Following AI conference, Cambridge University. Interdisciplinary: AI safety, AI governance, legal theory

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05244 2024-07-09 cs.AI cs.CL 54%

Some Issues in Predictive Ethics Modeling: An Annotated Contrast Set of "Moral Stories"

Ben Fitzgerald

专题命中 AI治理与伦理 :分类 cs.CL、cs.AI;safety(comments);AI safety(comments)

Comments This project was a runner-up to the Novel Research prize for the BlueDot Impact course AI Safety Fundamentals. View my contrast set as JSONL, the UI used to generate it, and Emelin et. al."s initial paper and code at https://github.com/bfitzgerald3132/MoralStoriesContrastSet

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05281 2026-08-28 cs.HC 版本更新 50%

Algorithmic Fairness Perceptions in the Global South: Evidence from Bangladesh on Ride-Sharing, Beauty Filters, and Large

全球南方的算法公平感知:来自孟加拉国关于网约车、美颜滤镜和大型语言模型的证据

Ahmed Abdal Shafi Rasel, Ahmed Mustafa Amlan, Tasmim Shajahan Mim

专题命中 AI治理与伦理 :trustworthy(abstract)

AI总结 该研究通过对孟加拉国199名民众的双语调查,揭示了情境影响算法公平感知、对算法保护需求普遍、美颜滤镜损害自我形象、识别与阐明LLM文化偏见存在差异等四个关键模式,指出仅结果导向的公平指标存在局限。

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14486 2026-08-25 cs.CV 版本更新 50%

Reduce the Artifacts Bias for More Generalizable AI-Generated Image Detection

降低艺术偏差以提高通用性的人工智能生成图像检测

Yiheng Li, Yang Yang, Wenhao Wang, Zichang Tan, Zecheng Lin, Li Gao, Zhen Lei

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS部门) Sangfor Technologies Inc.(Sangfor技术公司) China Mobile Financial Technology Co., Ltd.(中国移动金融科技有限公司) CAIR, HKSIS, Chinese Academy of Sciences(中国科学院CAIR、HKSIS部门) SCSE, FIE, M.U.S.T, Macau, China(澳门SCSE、FIE、M.U.S.T部门) Vast Intelligence Lab, Sydney, Australia(悉尼澳大利亚Vast Intelligence Lab)

专题命中 AI治理与伦理 :alignment(abstract)

AI总结 本文提出SEF框架,通过分离专家融合减少领域干扰,提升AI生成图像检测的通用性与鲁棒性,实验表明在13个基准上表现优异。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20413 2026-08-24 q-bio.GN 新提交 50%

BioFirewall: A genome-writing-native governance layer for design-stage biosecurity screening of agentic AI

BioFirewall:面向智能体AI设计阶段生物安全筛查的基因组写入原生治理层

Anees Ahmed Mahaboob Ali, Radhakrishnan Delhibabu, Everette Jacob Remington Nelson

专题命中 AI治理与伦理 :prompt injection(abstract)

AI总结 针对智能体AI基因组编辑设计阶段的生物安全管控缺口,提出BioFirewall中间件,其在多维度筛查中表现优于大模型,能有效抵御攻击且假拒绝率低,已开源并附带可复现基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18637 2026-08-21 cs.IR 版本更新 50%

PILOT Technical Report

PILOT技术报告

Jiuning Lin, Ruiquan Lan, Xiaodong Zhu, Bin Zhang, Chengyu Lai, Chuxin Chen, Dimin Wang, Han Zhu, Hongtao Cheng, Jialin Zhu, Lingqing Zhang, Shuai Zhong, Tao Wang, Weipeng Huang, Yinjiang Cai, Yinnan Song, Yuan Liu, Zhibo Xiao, Zhixin Ma, Zihong Huang

专题命中 AI治理与伦理 :safety(abstract)

AI总结 该研究针对现有推荐系统优化智能体方法的反应式缺陷,提出PILOT框架,通过三类角色构建控制循环,在淘宝平台对比ROAM验证,实现多项指标提升且搜索效率显著提高,无需人工干预。

Comments Technical Report, 42 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04306 2026-08-18 cs.MA 版本更新 50%

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems

组织控制层:LLM代理系统执行边界的治理基础设施

Tianyu Shi, Yang Mo, Yiou Liu, Zhuonan Hao, Yin Wang, Wenzhuo Hu, Nan Yu, Meng Zhou, Jiangbo Yu

专题命中 AI治理与伦理 :safety(abstract)

AI总结 针对LLM代理在执行边界产生的动作治理问题,提出组织控制层(OCL),通过策略执行和升级机制拦截生成动作,在不修改底层LLM生成器的情况下将不安全执行从88%降至接近零,同时将有效成功率从12%提升至96%。

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13969 2026-08-17 cs.CV 新提交 50%

PPOM: Marginalizing Patch-Grid Phase for CLIP-Based Generalizable Vision-Language Prompt Tuning

PPOM:用于基于CLIP的通用视觉-语言提示调优的补丁网格相位边缘化

Liang Wang, Haoyang Li, Chao Wang, Guodong Long, Jing Jiang, Yan Peng

机构 * Shanghai University(上海大学) University of Technology Sydney(悉尼科技大学)

专题命中 AI治理与伦理 :alignment(abstract)

AI总结 本研究针对基于CLIP的视觉-语言提示调优对补丁网格对齐敏感的问题,提出无训练的PPOM算子,通过边缘化相位偏移提升宿主性能,无需重新训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04341 2026-08-11 eess.SY cs.SY 版本更新 50%

Reinforcement Learning for Freeway Lane-Change Regulation via Connected Vehicles

基于网联车辆的高速公路车道变更调控强化学习方法

Ke Sun, Huan Yu

专题命中 AI治理与伦理 :safety(abstract)

AI总结 针对低自动驾驶渗透率下车道变更控制适用性受限的问题,提出基于多智能体强化学习的网联车辆高速公路车道变更调控框架,可提升交通效率并保障安全。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06829 2026-08-10 cs.SE 新提交 50%

How Reasoning Shapes Social Bias in LLM-Generated Code?

推理如何塑造大语言模型生成代码中的社会偏见?

Weifeng Sun, Jieke Shi, Zhou Yang, Yuchen Chen, Hongyan Li, Meng Yan, David Lo

专题命中 AI治理与伦理 :trustworthy(abstract)

AI总结 该研究首次系统探究基于推理的代码生成中的社会偏见,发现推理可降低偏见但会影响代码质量,进而提出ProbeDebias框架,有效检测并缓解代码偏见,同时保持较高质量。

Comments ASE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02827 2026-08-05 cs.MA 新提交 50%

Emergence of Biased Consensus in Multi-Agent LLM Debates

多智能体LLM辩论中偏向性共识的涌现

Maya Okawa

专题命中 AI治理与伦理 :safety(abstract)

AI总结 该研究针对多智能体LLM辩论中集体偏向性共识的涌现问题,提出物理启发的社会动力学分析框架,经实验验证其相变规律,发现智能体异质性可抑制该现象,且见解可推广至投资等现实决策任务。

Comments Accepted at the 43rd International Conference on Machine Learning (ICML 2026). 23 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01661 2026-08-04 cs.CV 新提交 50%

FairForensics: Seeing Expressions and Parsing Demographics via Vision-Language Modeling for Generalizable Fair Deepfake Detection

FairForensics:基于视觉-语言建模的通用公平深度伪造检测——表情感知与人口统计解析

Yaning Zhang, Jiao Wu, Zan Gao, Linlin Shen

机构 * Shenzhen University(深圳大学) Shandong Artificial Intelligence Institute(山东省人工智能研究院) Qilu University of Technology (Shandong Academy of Sciences)(齐鲁工业大学(山东省科学院)) Tianjin University of Technology(天津理工大学)

专题命中 AI治理与伦理 :alignment(abstract)

AI总结 本文构建人口统计平衡的深度伪造检测基准,提出FairForensics视觉-语言模型,通过表情感知与人口统计正则化实现通用公平检测,在基准上达到泛化与公平性的SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03697 2026-08-04 eess.AS 版本更新 50%

Improving ASR Fairness for Cleft Lip and Palate Speech: A Study on Severity-Aware Data Mixing

改善唇腭裂语音的自动语音识别公平性:一项关于严重程度感知数据混合的研究

Susmita Bhattacharjee, Jagabandhu Mishra, H. S. Shekhawat, Ravi Jasuja, S. R. Mahadeva Prasanna

专题命中 AI治理与伦理 :alignment(abstract)

AI总结 该研究针对唇腭裂(CLP)语音的ASR公平性问题,提出严重程度感知的CLP与正常语音混合策略,在AIISH和NMCPC数据集上使GMM-HMM、Whisper等模型的词错误率显著降低,提升了ASR公平性。

Comments Submitted to Speech Communication

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29610 2026-08-03 cs.SE 新提交 50%

Educating the Agentic Engineer: Curricula, Collaboration, and Continuous Learning in the AI Era

培养智能体工程师:AI时代的课程、协作与持续学习

Mamdouh Alenezi

专题命中 AI治理与伦理 :alignment(abstract)

AI总结 针对AI时代软件与系统工程的转型需求,本文提出ACCEL框架,构建智能体工程师的教育架构,明确核心能力与实施路径,识别风险并主张开展系统性教育变革。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29279 2026-08-03 cs.SD 新提交 50%

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition

ParaASR:面向快速长上下文基于大语言模型的语音识别的多令牌预测

Qingjian Lin, Yuxin Li, Haoyang Zhang, Jun Chen, Yechang Huang, Feng Tian, Xie Li, Xiangyu Tony Zhang, Daijiao Liu, Yuxin Zhang, Jinglan Gong, Bo Zhao, Fei Tian, Xuerui Yang, Gang Yu, Xiangyu Zhang, Daxin Jiang

机构 * StepFun(阶跃星辰) NTU(南洋理工大学) PKU(北京大学) UNSW(新南威尔士大学) SJTU(上海交通大学) USTC(中国科学技术大学)

专题命中 AI治理与伦理 :safety(abstract)

AI总结 ParaASR是一款基于大语言模型的ASR系统,通过多令牌预测技术,在保持低错误率、低延迟的同时,支持32K长上下文及30分钟音频的快速转录,兼顾识别质量、速度与上下文长度。

Comments 14 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏