arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2510.19002 2025-10-23 cs.GT cs.LG econ.TH math.OC 57%

Impartial Selection with Predictions

Javier Cembrano, Felix Fischer, Max Klimm

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20228 2025-10-23 cs.CR cs.CL 57%

Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID

Xia Han, Qi Li, Jianbing Ni, Mohammad Zulkernine

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) School of Computing(计算机学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted by TrustCom2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18425 2025-10-22 cs.AI 57%

Automated urban waterlogging assessment and early warning through a mixture of foundation models

Chenxu Zhang, Fuxiang Huang, Lei Zhang

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Submitted to Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18417 2025-10-22 cs.NI cs.AI 57%

On AI Verification in Open RAN

Rahul Soundrarajan, Claudio Fiandrino, Michele Polese, Salvatore D'Oro, Leonardo Bonati, Tommaso Melodia

机构 * Tejas Networks IMDEA Networks Institute Northeastern University

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06239 2025-10-22 cs.AI 57%

Proof2Silicon: Prompt Repair for Verified Code and Hardware Generation via Reinforcement Learning

Manvi Jha, Jiaxin Wan, Deming Chen

机构 * Electrical and Computer Engineering(电气与计算机工程系) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11690 2025-10-22 cs.LG cs.CV 57%

The Impact of Coreset Selection on Spurious Correlations and Group Robustness

Amaya Dharmasiri, William Yang, Polina Kirichenko, Lydia Liu, Olga Russakovsky

机构 * Princeton University(普林斯顿大学) FAIR at Meta(Meta 的 FAIR 实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 10 pages, 9 additional pages for Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17685 2025-10-21 cs.CV cs.AI 57%

Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning

Min Cao, Xinyu Zhou, Ding Jiang, Bo Du, Mang Ye, Min Zhang

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) School of Computer Science, Wuhan University(武汉大学计算机学院) Key Laboratory of New Generation Artificial Intelligence Technology & Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新 generation 人工智能技术及交叉应用重点实验室(东南大学),教育部,中国)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Final version published in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). Xplore link: https://ieeexplore.ieee.org/document/11199360

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17491 2025-10-21 cs.CL 57%

Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents

Yihong Tang, Kehai Chen, Liang Yue, Jinxin Fan, Caishen Zhou, Xiaoguang Li, Yuyang Zhang, Mingming Zhao, Shixiong Kai, Kaiyang Guo, Xingshan Zeng, Wenjing Cun, Lifeng Shang, Min Zhang

机构 * School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen, China(计算机科学与技术学院,哈尔滨工业大学,深圳,中国) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17263 2025-10-21 cs.CL 57%

TaxoAlign: Scholarly Taxonomy Generation Using Language Models

Avishek Lahiri, Yufang Hou, Debarshi Kumar Sanyal

机构 * Indian Association for the Cultivation of Science(印度科学培养协会) IT:U Interdisciplinary Transformation University Austria(奥地利 interdisciplinary Transformation University)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments This paper has been accepted at the EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17172 2025-10-21 cs.AI 57%

Combining ECG Foundation Model and XGBoost to Predict In-Hospital Malignant Ventricular Arrhythmias in AMI Patients

Shun Huang, Wenlu Xing, Shijia Geng, Hailong Wang, Guangkun Nie, Gongzheng Tang, Chenyang He, Shenda Hong

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17146 2025-10-21 cs.AI cs.CE 57%

Physics-Informed Large Language Models for HVAC Anomaly Detection with Autonomous Rule Generation

Subin Lin, Chuanbo Hua

机构 * Department of the Built Environment(环境学院) National University of Singapore(新加坡国立大学) InnoCore PRISM-AI KAIST(韩国科学技术院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments NeurIPS 2025 Workshop of UrbanAI (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03095 2025-10-21 cs.NI cs.AI cs.MA 57%

Evolution of AI Agent Registry Solutions: Centralized, Enterprise, and Distributed Approaches

Aditi Singh, Abul Ehtesham, Mahesh Lambe, Jared James Grogan, Abhishek Singh, Saket Kumar, Luca Muscariello, Vijoy Pandey, Guillaume Sauvage De Saint Marc, Pradyumna Chari, Ramesh Raskar

机构 * Cleveland State University(克利夫兰州立大学) Kent State University(肯特州立大学) Independent Researcher(独立研究者) Massachusetts Institute of Technology(麻省理工学院) Northeastern University(东北大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22904 2025-10-21 cs.HC cs.AI 57%

SketchMind: A Multi-Agent Cognitive Framework for Assessing Student-Drawn Scientific Sketches

Ehsan Latif, Zirak Khan, Xiaoming Zhai

机构 * AI4STEM Education Center University of Georgia(AI4STEM教育中心乔治亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Submitted to NeurIPS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02921 2025-10-21 cs.CL 57%

A Controllable Examination for Long-Context Language Models

Yijun Yang, Zeyu Huang, Wenhao Zhu, Zihan Qiu, Fei Yuan, Jeff Z. Pan, Ivan Titov

机构 * University of Edinburgh(爱丁堡大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Nanjing University(南京大学) Alibaba Group(阿里巴巴集团) University of Amsterdam(阿姆斯特丹大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments NeurIPS 2025 Dataset and Benchmark Track Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16652 2025-10-21 stat.ML cs.LG 57%

ARCO-BO: Adaptive Resource-aware COllaborative Bayesian Optimization for Heterogeneous Multi-Agent Design

Zihan Wang, Yi-Ping Chen, Tuba Dolar, Wei Chen

机构 * Department of Mechanical Engineering, Northwestern University(机械工程系,西北大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16611 2025-10-21 cs.CV cs.AI 57%

A Deep Learning Framework for Real-Time Image Processing in Medical Diagnostics: Enhancing Accuracy and Speed in Clinical Applications

Melika Filvantorkaman, Maral Filvan Torkaman

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 20 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16156 2025-10-21 eess.AS cs.AI cs.MM 57%

AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning

Yueqian Lin, Zhengmian Hu, Jayakumar Subramanian, Qinsi Wang, Nikos Vlassis, Hai "Helen" Li, Yiran Chen

机构 * Duke University(杜克大学) Adobe Research(Adobe研究)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted to the IEEE ASRU 2025 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16095 2025-10-21 cs.AI 57%

Reliability of Large Language Model Generated Clinical Reasoning in Assisted Reproductive Technology: Blinded Comparative Evaluation Study

Dou Liu, Ying Long, Sophia Zuoqiu, Di Liu, Kang Li, Yiting Lin, Hanyi Liu, Rong Yin, Tian Tang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11582 2025-10-21 cs.CL 57%

Subjective Evaluation Profile Analysis of Science Fiction Short Stories and its Critical-Theoretical Significance

Kazuyoshi Otsuka

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :RLHF(abstract);分类 cs.CL

Comments 38 pages. Manuscript submitted for review to the Journal of Computational Literary Studies (JCLS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10833 2025-10-21 cs.LG 57%

MergeBench: A Benchmark for Merging Domain-Specialized LLMs

Yifei He, Siqi Zeng, Yuzheng Hu, Rui Yang, Tong Zhang, Han Zhao

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments NeurIPS 2025 Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08750 2025-10-21 cs.CL 57%

HCR-Reasoner: Synergizing Large Language Models and Theory for Human-like Causal Reasoning

Yanxi Zhang, Xin Cong, Zhong Zhang, Xiao Liu, Dongyan Zhao, Yesai Wu

机构 * Center for Data Science, AAIS, Peking University(数据科学中心,AAIS,北京大学) Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学) Tsinghua University(清华大学) State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15739 2025-10-20 cs.AI cs.MA 57%

AURA: An Agent Autonomy Risk Assessment Framework

Lorenzo Satta Chiris, Ayush Mishra

机构 * University of Exeter(埃克塞特大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 10 pages, 2 figures. Submitted for open-access preprint on arXiv. Based on the AAMAS 2026 paper template

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15682 2025-10-20 cs.IR cs.CL 57%

SQuAI: Scientific Question-Answering with Multi-Agent Retrieval-Augmented Generation

Ines Besrour, Jingbo He, Tobias Schreieder, Michael Färber

机构 * TU Dresden(德累斯顿理工大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Accepted at CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15513 2025-10-20 cs.CL 57%

Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?

Ashutosh Bajpai, Tanmoy Chakraborty

机构 * Indian Institute of Technology Delhi(印度理工学院德里分校) MongoDB, Inc.(MongoDB公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments EMNLP Main Long Paper 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15422 2025-10-20 stat.ML cs.LG 57%

Information Theory in Open-world Machine Learning Foundations, Frameworks, and Future Direction

Lin Wang

机构 * Shenzhen Key Laboratory of Neuropsychiatric Modulation(深圳心理行为调控重点实验室) Shenzhen-Hong Kong Institute of Brain Science(深圳-香港脑科学研究院) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Chinese Academy of Sciences(中国科学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15306 2025-10-20 cs.AI 57%

WebGen-V Bench: Structured Representation for Enhancing Visual Design in LLM-based Web Generation and Evaluation

Kuang-Da Wang, Zhao Wang, Yotaro Shimose, Wei-Yao Wang, Shingo Takamatsu

机构 * National Yang Ming Chiao Tung University, Sony Group Corporation(National Yang Ming Chiao Tung University, Sony集团) Sony Group Corporation(Sony集团)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15106 2025-10-20 cs.CR cs.LG 57%

PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models

Issam Seddik, Sami Souihi, Mohamed Tamaazousti, Sara Tucci Piergiovanni

机构 * Université Paris-Saclay(巴黎-萨克雷大学) CEA LIST(法国原子能委员会列表中心) Palaiseau, France(法国Palaiseau)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 10 pages, 6 figures, 1 table. Accepted for presentation at FLLM 2025 (Vienna, Nov 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07176 2025-10-20 cs.MA cs.AI 57%

Internet of Agents: Fundamentals, Applications, and Challenges

Yuntao Wang, Shaolong Guo, Yanghe Pan, Zhou Su, Fahao Chen, Tom H. Luan, Peng Li, Jiawen Kang, Dusit Niyato

机构 * School of Cyber Science and Engineering, Xi'an Jiaotong University(网络安全科学与工程学院,西安交通大学) School of Artificial Intelligence, Shandong University(人工智能学院,山东大学) School of Automation, Guangdong University of Technology(自动化学院,广东技术大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 25 pages,10 figures, 10 tables. Accepted by IEEE TCCN in Oct. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15176 2025-10-20 cs.CV cs.AI 57%

Methods and Trends in Detecting AI-Generated Images: A Comprehensive Review

Arpan Mahara, Naphtali Rishe

机构 * Knight Foundation School of Computing and Information Sciences, Florida International University(骑士基金会计算与信息科学学院,佛罗里达国际大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 34 pages, 4 Figures, 10 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14525 2025-10-17 cs.CV cs.AI 57%

Real-Time Surgical Instrument Defect Detection via Non-Destructive Testing

Qurrat Ul Ain, Atif Aftab Ahmed Jilani, Zunaira Shafqat, Nigar Azhar Butt

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏