arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2509.09871 2025-09-22 cs.CL cs.AI 62%

Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case

Bastián González-Bustamante, Nando Verelst, Carla Cisternas

机构 * Universidad Diego Portales(迪亚戈·波特莱斯大学) Leiden University(莱顿大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Working paper: 18 pages, 4 tables, 2 figures

Journal ref Empiria Lab Method Series (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13759 2025-09-22 cs.LG cs.AI 62%

Discrete Diffusion in Large Language and Multimodal Models: A Survey

Runpeng Yu, Qi Li, Xinchao Wang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13951 2025-09-22 cs.CL cs.AI 62%

A Layered Multi-Expert Framework for Long-Context Mental Health Assessments

Jinwen Tang, Qiming Guo, Wenbo Sun, Yi Shang

机构 * Texas A\&M University-Corpus Christi(德克萨斯A&M大学-科珀斯克里斯蒂) Delft University of Technology(代尔夫特理工大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Journal ref Proc. 2025 IEEE Conference on Artificial Intelligence (CAI), Santa Clara, CA, USA, 2025, pp. 435-440

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14930 2025-09-19 cs.CL cs.AI 62%

Cross-Modal Knowledge Distillation for Speech Large Language Models

Enzhi Wang, Qicheng Li, Zhiyuan Tang, Yuhang Jia

机构 * TMCC, College of Computer Science, Nankai University, Tianjin, China(TMCC,计算机科学学院,南开大学,天津,中国) Tencent Ethereal Audio Lab, Tencent Corporation, Shenzhen, China(腾讯虚实音频实验室,腾讯公司,深圳,中国)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07054 2025-09-19 cs.AI cs.LG stat.ME 62%

Statistical Methods in Generative AI

Edgar Dobriban

机构 * Edgar Dobriban 1(Edgar Dobriban)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Invited review paper for Annual Review of Statistics and Its Application. Feedback welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20020 2025-09-19 cs.LG cs.AI 62%

Modular Machine Learning: An Indispensable Path towards New-Generation Large Language Models

Xin Wang, Haoyang Li, Haibo Chen, Zeyang Zhang, Wenwu Zhu

机构 * Department of Computer Science and Technology, BNRist, Tsinghua University(计算机科学与技术系,BNRist,清华大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 20 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10970 2025-09-18 cs.LG cs.AI 62%

The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in Large Language Models

Joshua Au Yeung, Jacopo Dalmasso, Luca Foschini, Richard JB Dobson, Zeljko Kraljevic

机构 * King’s College Hospital(国王学院医院) Nuraxi AI Dev and Doc: AI for Healthcare(Nuraxi AI 人工智能医疗) Sage Bionetworks(Sage 生物网络) University College London(伦敦大学学院) King’s College London(国王学院伦敦)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12740 2025-09-17 cs.RO cs.AI cs.ET cs.LG cs.SY eess.SY 62%

Deep Generative and Discriminative Digital Twin endowed with Variational Autoencoder for Unsupervised Predictive Thermal Condition Monitoring of Physical Robots in Industry 6.0 and Society 6.0

Eric Guiffo Kaigom

机构 * Department of Computer Science \& Engineering, Frankfurt University of Applied Sciences, Frankfurt a.M., Germany (e-mail: ).

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments $©$ 2025 the authors. This work has been accepted to the to the 10th IFAC Symposium on Mechatronic Systems & 14th IFAC Symposium on Robotics July 15-18, 2025 || Paris, France for publication under a Creative Commons Licence CC-BY-NC-ND

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12259 2025-09-17 cs.LG cs.AI quant-ph 62%

Quantum-Inspired Stacked Integrated Concept Graph Model (QISICGM) for Diabetes Risk Prediction

Kenneth G. Young

机构 * II (September 12, 2025)(II)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 3 figures, includes performance tables and visualizations. Proposes a Quantum-Inspired Stacked Integrated Concept Graph Model (QISICGM) that integrates phase feature mapping, self-improving concept graphs, and neighborhood sequence modeling within a stacked ensemble. Demonstrates improved F1 and AUC on an augmented PIMA Diabetes dataset with efficient CPU inference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11595 2025-09-16 cs.AI cs.CE cs.CR cs.LG cs.MA 62%

AMLNet: A Knowledge-Based Multi-Agent Framework to Generate and Detect Realistic Money Laundering Transactions

Sabin Huda, Ernest Foo, Zahra Jadidi, MA Hakim Newton, Abdul Sattar

机构 * School of Information and Communication Technology, Griffith University, QLD Australia(信息与通信技术学院,格里菲斯大学,昆士兰州澳大利亚) School of Information and Physical Sciences, The University of Newcastle, NSW Australia(信息与物理科学学院,新castle大学,新南威尔士州澳大利亚)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11376 2025-09-16 cs.LG cs.AI cs.CE 62%

Intelligent Reservoir Decision Support: An Integrated Framework Combining Large Language Models, Advanced Prompt Engineering, and Multimodal Data Fusion for Real-Time Petroleum Operations

Seyed Kourosh Mahjour, Seyed Saman Mahjour

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10802 2025-09-16 q-fin.RM cs.CL cs.LG q-fin.CP 62%

Why Bonds Fail Differently? Explainable Multimodal Learning for Multi-Class Default Prediction

Yi Lu, Aifan Ling, Chaoqun Wang, Yaxin Xu

机构 * School of Economics and Finance, Shanghai International Studies University(经济金融学院,上海国际问题研究大学) School of AI and Advanced Computing, Xi’an Jiaotong-Liverpool University(人工智能与先进计算学院,西安交通大学利物浦大学) School of Foreign Studies, Shanghai University of Finance and Economics(外国语言学院,上海金融学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09593 2025-09-12 cs.CL cs.AI 62%

Fluent but Unfeeling: The Emotional Blind Spots of Language Models

Bangzhao Shu, Isha Joshi, Melissa Karnaze, Anh C. Pham, Ishita Kakkar, Sindhu Kothe, Arpine Hovasapian, Mai ElSherief

机构 * Northeastern University(东北大学) UC San Diego(加州大学圣地亚哥分校) University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Camera-ready version for ICWSM 2026. First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04482 2025-09-09 cs.CL cs.AI 62%

Energy Landscapes Enable Reliable Abstention in Retrieval-Augmented Large Language Models for Healthcare

Ravi Shankar, Sheng Wong, Lin Li, Magdalena Bachmann, Alex Silverthorne, Beth Albert, Gabriel Davis Jones

机构 * Oxford Digital Health Labs(牛津数字健康实验室) Nuffield Department of Women’s and Reproductive Health(妇女与生殖健康尼富尔德部门) University of Oxford(牛津大学) OATML Department of Computer Science(计算机科学系)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05380 2025-09-09 cs.CY cs.AI cs.CR cs.RO 62%

Cumplimiento del Reglamento (UE) 2024/1689 en robótica y sistemas autónomos: una revisión sistemática de la literatura

Yoana Pita Lorenzo

机构 * Universidad de León(莱昂大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

Comments in Spanish language

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04752 2025-09-08 cs.HC cs.AI cs.LG 62%

SePA: A Search-enhanced Predictive Agent for Personalized Health Coaching

Melik Ozolcer, Sang Won Bae

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted at IEEE-EMBS International Conference on Biomedical and Health Informatics (BHI'25). 7 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05528 2025-09-08 cs.AI cs.CL 62%

Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment

Jiahuan Pei, Fanghua Ye, Xin Sun, Wentao Deng, Koen Hindriks, Junxiao Wang

机构 * Vrije University of Amsterdam(阿姆斯特丹自由大学) University College London(伦敦大学学院) University of Amsterdam(阿姆斯特丹大学) National Institute of Informatics(日本信息处理学会) Shandong University(山东大学) Guangzhou University(广州大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 pages, accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23903 2025-09-08 cs.AI cs.LG 62%

Neural Network Verification with PyRAT

Augustin Lemesle, Julien Lehmann, Tristan Le Gall

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04504 2025-09-08 cs.CL cs.AI 62%

Behavioral Fingerprinting of Large Language Models

Zehua Pei, Hui-Ling Zhen, Ying Zhang, Zhiyuan Yang, Xing Li, Xianzhi Yu, Mingxuan Yuan, Bei Yu

机构 * The Chinese University of Hong Kong(香港中文大学) Noah’s Ark Lab, Huawei(华为诺亚实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Submitted to 1st Open Conference on AI Agents for Science (agents4science 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04499 2025-09-08 cs.CL cs.AI 62%

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou, Kung-Hsiang Huang, Yixin Mao, Chien-Sheng Wu

机构 * Salesforce AI Research(Salesforce AI研究)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: text overlap with arXiv:2410.22349

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03809 2025-09-05 cs.CL cs.AI 62%

Align-then-Slide: A complete evaluation framework for Ultra-Long Document-Level Machine Translation

Jiaxin Guo, Daimeng Wei, Yuanchang Luo, Xiaoyu Chen, Zhanglin Wu, Huan Yang, Hengchao Shang, Zongyao Li, Zhiqiang Rao, Jinlong Yang, Hao Yang

机构 * Huawei Translation Services Center(华为翻译服务中心)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments under preview

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03529 2025-09-05 cs.CL cs.AI eess.AS 62%

Multimodal Proposal for an AI-Based Tool to Increase Cross-Assessment of Messages

Alejandro Álvarez Castro, Joaquín Ordieres-Meré

机构 * AI master(人工智能硕士) Universidad Politécnica de Madrid(马德里理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Presented at NLMLT2025 (https://airccse.org/csit/V15N16.html), 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03329 2025-09-04 cs.CY cs.CL 62%

SESGO: Spanish Evaluation of Stereotypical Generative Outputs

Melissa Robles, Catalina Bernal, Denniss Raigoso, Mateo Dulce Rubio

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03169 2025-09-04 cs.LG cs.AI 62%

Rashomon in the Streets: Explanation Ambiguity in Scene Understanding

Helge Spieker, Jørn Eirik Betten, Arnaud Gotlieb, Nadjib Lazaar, Nassim Belmecheri

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments AAAI 2025 Fall Symposium: AI Trustworthiness and Risk Assessment for Challenged Contexts (ATRACC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02855 2025-09-04 cs.CL cs.CY 62%

IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations

Hyunji Nam, Lucia Langlois, James Malamut, Mei Tan, Dorottya Demszky

机构 * Stanford University(斯坦福大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.CY

Comments 10 pages, 9 pages for appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21376 2025-09-04 cs.AI cs.CL 62%

AHELM: A Holistic Evaluation of Audio-Language Models

Tony Lee, Haoqin Tu, Chi Heem Wong, Zijun Wang, Siwei Yang, Yifan Mai, Yuyin Zhou, Cihang Xie, Percy Liang

机构 * Stanford University(斯坦福大学) University of California, Santa Cruz(加州大学圣克ruz分校) Hitachi America, Ltd.(日立美国有限公司)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02350 2025-09-03 cs.CL cs.AI 62%

Implicit Reasoning in Large Language Models: A Comprehensive Survey

Jindong Li, Yali Fu, Li Fan, Jiahong Liu, Yao Shu, Chengwei Qin, Menglin Yang, Irwin King, Rex Ying

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Jilin University(吉林大学) The Chinese University of Hong Kong(香港中文大学) Yale University(耶鲁大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00414 2025-09-03 cs.CL cs.AI cs.IR 62%

MedSEBA: Synthesizing Evidence-Based Answers Grounded in Evolving Medical Literature

Juraj Vladika, Florian Matthes

机构 * Technical University of Munich Department of Computer Science(慕尼黑技术大学计算机科学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted to CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07407 2025-09-03 cs.AI cs.CL cs.MA 62%

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems

Jinyuan Fang, Yanwen Peng, Xi Zhang, Yingxu Wang, Xinhao Yi, Guibin Zhang, Yi Xu, Bin Wu, Siwei Liu, Zihao Li, Zhaochun Ren, Nikos Aletras, Xi Wang, Han Zhou, Zaiqiao Meng

机构 * University of Glasgow(格拉斯哥大学) University of Sheffield(谢菲尔德大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) National University of Singapore(新加坡国家大学) University of Cambridge(剑桥大学) University College London(伦敦大学学院) University of Aberdeen(阿伯丁大学) Leiden University(莱顿大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments Github Repo: https://github.com/EvoAgentX/Awesome-Self-Evolving-Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09318 2025-09-03 cs.CL cs.AI 62%

Benchmarking LLMs for Mimicking Child-Caregiver Language in Interaction

Jing Liu, Abdellah Fourtassi

机构 * ENS, PSL Research University, EHESS, CNRS, France(巴黎社会科学高等学院、巴黎综合理工研究大学、高等社会科学研究院、法国国家科学研究中心)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏