arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9400 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9400 篇

2602.13455 2026-02-17 cs.CL cs.AI cs.HC 76%

Using Machine Learning to Enhance the Detection of Obfuscated Abusive Words in Swahili: A Focus on Child Safety

利用机器学习增强斯瓦希里语中隐晦侮辱性词汇的检测:聚焦儿童安全

Phyllis Nabangi, Abdul-Jalil Zakaria, Jema David Ndibwile

专题命中 安全评测 :safety(title);分类 cs.CL、cs.AI

AI总结 本研究利用机器学习方法提升斯瓦希里语中隐晦侮辱性词汇的检测能力,旨在提高儿童网络环境的安全性。

Comments Accepted at the Second IJCAI AI for Good Symposium in Africa, hosted by Deep Learning Indaba, 7 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07414 2026-02-10 cs.AI cs.CL 76%

Can LLMs Truly Embody Human Personality? Analyzing AI and Human Behavior Alignment in Dispute Resolution

LLMs真的能体现人类个性吗?分析AI与人类行为在纠纷解决中的对齐

Deuksin Kwon, Kaleen Shrestha, Bin Han, Spencer Lin, James Hale, Jonathan Gratch, Maja Matarić, Gale M. Lucas

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

AI总结 本文研究LLMs在模拟人类人格驱动的冲突行为方面的有效性,通过评估框架和数据集方法,揭示LLMs在纠纷解决中与人类行为的显著差异。

Comments AAAI 2026 (Special Track: AISI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00078 2026-02-03 cs.CY cs.AI 76%

Standards for trustworthy AI in the European Union: technical rationale, structural challenges, and an implementation path

欧盟可信人工智能标准:技术依据、结构性挑战及实施路径

Piercosma Bisconti, Marcello Galisai

机构 * Department of Computer, Control and Management Engineering(计算机、控制与管理工程系) Sapienza University of Rome(罗马萨皮恩扎大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.CY

AI总结 本文提出欧盟人工智能标准化的实施路径,强调通过分层方法和风险管理体系,实现技术标准与法律义务的对接。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21900 2026-02-03 cs.CV cs.AI cs.CY cs.MM 76%

TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention

TraceRouter: 通过路径级干预实现大型基础模型的鲁棒安全性

Chuancheng Shi, Shangze Li, Wenjun Lu, Wenhua Wu, Cong Wang, Zifeng Cheng, Fei Shen, Tat-Seng Chua

机构 * The University of Sydney, Sydney, Australia(悉尼大学) Nanjing University of Science(南京理工大学) Nanjing University, Nanjing, China(南京大学) National University of Singapore, Singapore, Singapore(新加坡国立大学)

专题命中 安全评测 :safety(title);分类 cs.AI、cs.CY

AI总结 TraceRouter通过路径级干预有效切断有害语义的因果传播,提升大型基础模型的对抗鲁棒性与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23845 2026-01-26 cs.CL cs.AI 76%

CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection

CRADLE Bench: 一种多维心理健康危机和安全风险检测的临床标注基准

Grace Byun, Rebecca Lipschutz, Sean T. Minton, Abigail Lott, Jinho D. Choi

机构 * Emory University, Department of Computer Science(埃默里大学计算机科学系) Emory University, Department of Psychiatry and Behavioral Sciences(埃默里大学精神病学与行为科学系)

专题命中 安全评测 :safety(title);分类 cs.CL、cs.AI

AI总结 CRADLE Bench通过多维危机检测基准,结合临床标准和时间标签,提升语言模型对心理健康危机和安全风险的识别能力。

Journal ref EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10645 2025-12-09 cs.LG cs.AI 76%

Trustworthy Retrosynthesis: Eliminating Hallucinations with a Diverse Ensemble of Reaction Scorers

可信的逆合成:通过多样化的反应评分器消除幻觉

Michal Sadowski, Tadija Radusinović, Maria Wyrzykowska, Lukasz Sztukiewicz, Jan Rzymkowski, Paweł Włodarczyk-Pruszyński, Mikołaj Sacha, Piotr Kozakowski, Ruard van Workum, Stanislaw Kamil Jastrzebski

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

AI总结 RetroTrim通过多样化反应评分策略有效消除逆合成中的幻觉,实现高质量路径生成,解决药物类似物领域中的合成计划可靠性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04184 2025-11-07 cs.CL cs.AI 76%

Trustworthy LLM-Mediated Communication: Evaluating Information Fidelity in LLM as a Communicator (LAAC) Framework in Multiple Application Domains

Mohammed Musthafa Rafi, Adarsh Krishnamurthy, Aditya Balu

机构 * Iowa State University(爱荷华州立大学)

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI

Comments 10 pages, 4 figures. Submitted to IEEE DISTILL 2025 (co-located with IEEE TPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03665 2025-11-04 cs.LG cs.AI 76%

A DbC Inspired Neurosymbolic Layer for Trustworthy Agent Design

Claudiu Leoveanu-Condrei

机构 * ExtensityAI

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments 4 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08525 2025-10-31 cs.LG cs.AI 76%

A mathematical certification for positivity conditions in Neural Networks with applications to partial monotonicity and Trustworthy AI

Alejandro Polo-Molina, David Alfaya, Jose Portela

机构 * CDTI

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments 16 pages, 4 figures

Journal ref IEEE Transactions on Neural Networks and Learning Systems, Early Access, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09090 2025-10-13 cs.CY cs.AI 76%

AI and Human Oversight: A Risk-Based Framework for Alignment

Laxmiraju Kandikatla, Branislav Radeljic

专题命中 安全评测 :alignment(title);分类 cs.AI、cs.CY

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08608 2025-10-13 cs.CL cs.AI 76%

MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation

Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Agency for Science, Technology and Research, Singapore(新加坡科技研究局) Indian Institute of Technology Delhi(印度理工学院德里分校) Alibaba DAMO Academy(阿里巴巴达摩院) Microsoft Research Asia(微软亚洲研究院) Shanghai University of Finance and Economics(上海财经大学) Inner Mongolia University(内蒙古大学) Kyoto University(京都大学) Jiangxi Normal University(江西师范大学) Korea University(韩国大学) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08429 2025-10-10 cs.LG cs.AI stat.ML 76%

ClauseLens: Clause-Grounded, CVaR-Constrained Reinforcement Learning for Trustworthy Reinsurance Pricing

Stella C. Dong, James R. Finlay

机构 * Department of Applied Mathematics, University of California, Davis, CA, USA(加州大学戴维斯分校应用数学系) Wharton School of Business, University of Pennsylvania, Philadelphia, PA, USA(宾夕法尼亚大学沃顿商学院)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments Accepted for publication at the 6th ACM International Conference on AI in Finance (ICAIF 2025), Singapore. Author-accepted version (October 2025). 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14756 2025-10-10 cs.LG cs.AI 76%

LLINBO: Trustworthy LLM-in-the-Loop Bayesian Optimization

Chih-Yu Chang, Milad Azvar, Chinedum Okwudire, Raed Al Kontar

机构 * University of Michigan(密歇根大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19120 2025-09-24 cs.LG cs.AI cs.DC 76%

FedFiTS: Fitness-Selected, Slotted Client Scheduling for Trustworthy Federated Learning in Healthcare AI

Ferdinand Kahenga, Antoine Bagula, Sajal K. Das, Patrick Sello

机构 * Department of Computer Science University of the Western Cape(计算机科学系,西开普敦大学) Department of Computer Science Missouri University of Science and Technology(计算机科学系,密苏里科学与技术大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17764 2025-09-22 cs.CL cs.AI math.ST stat.TH 76%

BBScoreV2: Learning Time-Evolution and Latent Alignment from Stochastic Representation

Tianhao Zhang, Zhecheng Sheng, Zhexiao Lin, Chen Jiang, Dongyeop Kang

机构 * University of Minnesota, Twin Cities(明尼苏达大学,双城分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

Journal ref The 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12233 2025-09-17 cs.CR cs.AI cs.ET cs.LG cs.NI 76%

Towards Trustworthy Agentic IoEV: AI Agents for Explainable Cyberthreat Mitigation and State Analytics

Meryem Malak Dif, Mouhamed Amine Bouchiha, Abdelaziz Amara Korba, Yacine Ghamri-Doudane

机构 * L3i - La Rochelle University, La Rochelle, France(L3i - 拉罗谢尔大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments 10 pages, 7 figures, Accepted at LCN'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06902 2025-09-09 cs.CL cs.CR cs.DB cs.LG 76%

Proof-Carrying Numbers (PCN): A Protocol for Trustworthy Numeric Answers from LLMs via Claim Verification

Aivin V. Solatorio

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17114 2025-09-08 cs.CL cs.CV cs.LG cs.MM 76%

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language

Subrata Biswas, Mohammad Nur Hossain Khan, Bashima Islam

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24671 2025-09-03 cs.CL cs.AI 76%

Multiple LLM Agents Debate for Equitable Cultural Alignment

Dayeon Ki, Rachel Rudinger, Tianyi Zhou, Marine Carpuat

机构 * University of Maryland(马里兰大学)

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

Comments ACL 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22940 2025-08-05 cs.CL cs.AI 76%

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes

Rui Jiao, Yue Zhang, Jinku Li

机构 * School of Cyber Engineering, Xidian University(西安电子科技大学电子工程学院) School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院)

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17010 2025-07-24 cs.CR cs.AI cs.LG 76%

Towards Trustworthy AI: Secure Deepfake Detection using CNNs and Zero-Knowledge Proofs

H M Mohaimanul Islam, Huynh Q. N. Vo, Aditya Rane

机构 * School of Industrial Engineering and Management(工业工程与管理学院)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments Submitted for peer-review in TrustXR - 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13524 2025-07-21 cs.HC cs.AI cs.CY 76%

Humans learn to prefer trustworthy AI over human partners

Yaomin Jiang, Levin Brinkmann, Anne-Marie Nussberger, Ivan Soraperra, Jean-François Bonnefon, Iyad Rahwan

机构 * Toulouse School of Economics, Centre National de la Recherche Scientifique (TSM-R), Université Toulouse Capitole(图卢兹经济学院,法国国家科学研究中心(TSM-R),图卢兹大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10240 2025-07-15 cs.HC cs.AI cs.LG 76%

Visual Analytics for Explainable and Trustworthy Artificial Intelligence

Angelos Chatzimparmpas

机构 * Utrecht University(乌特勒支大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Journal ref IEEE CG&A 2025, vol. 45, pp. 100-111

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07576 2025-07-11 cs.AI cs.LG cs.LO 76%

On Trustworthy Rule-Based Models and Explanations

Mohamed Siala, Jordi Planes, Joao Marques-Silva

机构 * LAAS-CNRS, Université de Toulouse, CNRS, INSA Toulouse, France(法国图卢兹大学、CNRS、INSA图卢兹分校) Universitat de Lleida(莱里达大学) ICREA & University of Lleida(ICREA与莱里达大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01281 2025-07-03 cs.CL cs.AI 76%

Rethinking All Evidence: Enhancing Trustworthy Retrieval-Augmented Generation via Conflict-Driven Summarization

Juan Chen, Baolong Bi, Wei Zhang, Jingyan Sui, Xiaofei Zhu, Yuanzhuo Wang, Lingrui Mei, Shenghua Liu

机构 * University of Chinese Academy of Sciences(中国科学院大学) Chinese Academy of Sciences(中国科学院) National University of Defense Technology(国防科技大学) Chongqing University of Technology(重庆理工大学)

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02080 2025-06-12 cs.CV cs.CL cs.LG 76%

EMMA: Efficient Visual Alignment in Multi-Modal LLMs

Sara Ghazanfari, Alexandre Araujo, Prashanth Krishnamurthy, Siddharth Garg, Farshad Khorrami

机构 * Department of Electronic and Computer Engineering, New York University(电子与计算机工程系,纽约大学)

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00519 2025-06-04 cs.CL cs.AI 76%

CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention

Yuxi Sun, Aoqi Zuo, Wei Gao, Jing Ma

机构 * Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系) School of Mathematics and Statistics, The University of Melbourne(墨尔本大学数学与统计学学院) School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算与信息系统学院)

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI

Comments Accepted to Association for Computational Linguistics Findings (ACL) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02017 2025-06-02 cs.SI cs.AI cs.LG 76%

Multi-Domain Graph Foundation Models: Robust Knowledge Transfer via Topology Alignment

Shuo Wang, Bokui Wang, Zhixiang Shen, Boyan Deng, Zhao Kang

专题命中 安全评测 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16103 2025-05-23 cs.LG cs.AI 76%

Towards Trustworthy Keylogger detection: A Comprehensive Analysis of Ensemble Techniques and Feature Selections through Explainable AI

Monirul Islam Mahmud

机构 * Dept. of Computer & Information Science(计算机与信息科学系) Fordham University(福特汉姆大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13028 2025-05-21 cs.CR cs.AI cs.CL 76%

Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset

Sayon Palit, Daniel Woods

机构 * School of Informatics University of Edinburgh(信息学院爱丁堡大学)

专题命中 安全评测 :safety(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏