arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-20 至 2025-10-20 共收录 21 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 21 篇

2510.13023 2025-10-20 cs.LG physics.comp-ph 79%

Machine Learning-Based Ultrasonic Weld Characterization Using Hierarchical Wave Modeling and Diffusion-Driven Distribution Alignment

Joshua R. Tempelman, Adam J. Wachtor, Eric B. Flynn

机构 * Data Science Group, Los Alamos National Laboratory, Los Alamos NM, USA(数据科学组,洛斯阿拉莫斯国家实验室) Engineering Institute, Los Alamos National Laboratory, Los Alamos NM, USA(工程学院,洛斯阿拉莫斯国家实验室)

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

Comments 26 pages, 6 page appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14398 2025-10-20 cs.CL 79%

Lightweight Safety Guardrails Using Fine-tuned BERT Embeddings

Aaron Zheng, Mansi Rana, Andreas Stolcke

机构 * Uniphore & UC Berkeley(Uniphore与伯克利大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments To appear in Proceedings of COLING 2025

Journal ref Proc. 31st Intl. Conf. Computational Linguistics: Industry Track, COLING 2025, pp. 689-696

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15232 2025-10-20 cs.LG cs.CL 73%

FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain

Tiansheng Hu, Tongyan Hu, Liuyang Bai, Yilun Zhao, Arman Cohan, Chen Zhao

机构 * NYU Shanghai(纽约大学上海校区) National University of Singapore(新加坡国立大学) Yale University(耶鲁大学) Center for Data Science, New York University(纽约大学数据科学中心)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15866 2025-10-20 cs.CV cs.NE 67%

BiomedXPro: Prompt Optimization for Explainable Diagnosis with Biomedical Vision Language Models

Kaushitha Silva, Mansitha Eashwara, Sanduni Ubayasiri, Ruwan Tennakoon, Damayanthi Herath

机构 * University of Peradeniya(珀德尼亚大学) RMIT University(皇家墨尔本理工大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract)

Comments 10 Pages + 15 Supplementary Material Pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21603 2025-10-20 cs.CL cs.CY cs.LG 67%

Operationalizing Automated Essay Scoring: A Human-Aware Approach

Yenisel Plasencia-Calaña

机构 * Brigthlands Institute for a Smart Society(智能社会研究院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15768 2025-10-20 cs.CL cs.LG 62%

On Non-interactive Evaluation of Animal Communication Translators

Orr Paradise, David F. Gruber, Adam Tauman Kalai

机构 * EPFL(瑞士联邦理工学院) Project CETI(CETI项目) OpenAI

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15211 2025-10-20 cs.LG cs.AI 62%

ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning

Yongchan Kwon, Shang Zhu, Federico Bianchi, Kaitlyn Zhou, James Zou

机构 * Together AI Stanford University(斯坦福大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15007 2025-10-20 cs.CL cs.AI 62%

Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective

Zhiqiang Kou, Junyang Chen, Xin-Qiang Cai, Ming-Kun Xie, Biao Liu, Changwei Wang, Lei Feng, Yuheng Jia, Gang Niu, Masashi Sugiyama, Xin Geng

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) RIKEN Center for Advanced Intelligence Project (AIP)(RIKEN先进人工智能项目中心) School of Computer Science and Technology, Qilu University of Technology(齐鲁大学计算机科学与技术学院) The University of Tokyo(东京大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10444 2025-10-20 cs.CL cs.AI 62%

Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance

Jingyi Chen, Zhimeng Guo, Jiyun Chun, Pichao Wang, Andrew Perrault, Micha Elsner

机构 * Department of Linguistics, The Ohio State University, USA(语言学系,俄亥俄州立大学) Department of Computer Science and Engineering, The Ohio State University, USA(计算机科学与工程系,俄亥俄州立大学) Department of Information Sciences and Technology, Penn State University, USA(信息科学与技术系,宾夕法尼亚州立大学) Amazon, USA(亚马逊公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15739 2025-10-20 cs.AI cs.MA 57%

AURA: An Agent Autonomy Risk Assessment Framework

Lorenzo Satta Chiris, Ayush Mishra

机构 * University of Exeter(埃克塞特大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 10 pages, 2 figures. Submitted for open-access preprint on arXiv. Based on the AAMAS 2026 paper template

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15682 2025-10-20 cs.IR cs.CL 57%

SQuAI: Scientific Question-Answering with Multi-Agent Retrieval-Augmented Generation

Ines Besrour, Jingbo He, Tobias Schreieder, Michael Färber

机构 * TU Dresden(德累斯顿理工大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Accepted at CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15513 2025-10-20 cs.CL 57%

Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?

Ashutosh Bajpai, Tanmoy Chakraborty

机构 * Indian Institute of Technology Delhi(印度理工学院德里分校) MongoDB, Inc.(MongoDB公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments EMNLP Main Long Paper 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15422 2025-10-20 stat.ML cs.LG 57%

Information Theory in Open-world Machine Learning Foundations, Frameworks, and Future Direction

Lin Wang

机构 * Shenzhen Key Laboratory of Neuropsychiatric Modulation(深圳心理行为调控重点实验室) Shenzhen-Hong Kong Institute of Brain Science(深圳-香港脑科学研究院) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Chinese Academy of Sciences(中国科学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15306 2025-10-20 cs.AI 57%

WebGen-V Bench: Structured Representation for Enhancing Visual Design in LLM-based Web Generation and Evaluation

Kuang-Da Wang, Zhao Wang, Yotaro Shimose, Wei-Yao Wang, Shingo Takamatsu

机构 * National Yang Ming Chiao Tung University, Sony Group Corporation(National Yang Ming Chiao Tung University, Sony集团) Sony Group Corporation(Sony集团)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15106 2025-10-20 cs.CR cs.LG 57%

PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models

Issam Seddik, Sami Souihi, Mohamed Tamaazousti, Sara Tucci Piergiovanni

机构 * Université Paris-Saclay(巴黎-萨克雷大学) CEA LIST(法国原子能委员会列表中心) Palaiseau, France(法国Palaiseau)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 10 pages, 6 figures, 1 table. Accepted for presentation at FLLM 2025 (Vienna, Nov 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07176 2025-10-20 cs.MA cs.AI 57%

Internet of Agents: Fundamentals, Applications, and Challenges

Yuntao Wang, Shaolong Guo, Yanghe Pan, Zhou Su, Fahao Chen, Tom H. Luan, Peng Li, Jiawen Kang, Dusit Niyato

机构 * School of Cyber Science and Engineering, Xi'an Jiaotong University(网络安全科学与工程学院,西安交通大学) School of Artificial Intelligence, Shandong University(人工智能学院,山东大学) School of Automation, Guangdong University of Technology(自动化学院,广东技术大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 25 pages,10 figures, 10 tables. Accepted by IEEE TCCN in Oct. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15176 2025-10-20 cs.CV cs.AI 57%

Methods and Trends in Detecting AI-Generated Images: A Comprehensive Review

Arpan Mahara, Naphtali Rishe

机构 * Knight Foundation School of Computing and Information Sciences, Florida International University(骑士基金会计算与信息科学学院,佛罗里达国际大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 34 pages, 4 Figures, 10 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15729 2025-10-20 cs.IR 50%

FACE: A General Framework for Mapping Collaborative Filtering Embeddings into LLM Tokens

Chao Wang, Yixin Song, Jinhui Ye, Chuan Qin, Dazhong Shen, Lingfeng Liu, Xiang Wang, Yanyong Zhang

专题命中 安全评测 :alignment(abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15564 2025-10-20 cs.CV 50%

Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation

Xiaoming Zhu, Xu Huang, Qinghongbing Xie, Zhi Deng, Junsheng Yu, Yirui Guan, Zhongyuan Liu, Lin Zhu, Qijun Zhao, Ligang Liu, Long Zeng

机构 * Tsinghua University(清华大学) Tencent(腾讯) Southeast University(东南大学) University of Science and Technology of China(中国科学技术大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15397 2025-10-20 cond-mat.mtrl-sci 50%

Unravelling the Catalytic Activity of Dual-Metal Doped N6-Graphene for Sulfur Reduction via Machine Learning-Accelerated First-Principles Calculations

Sahil Kumar, Adithya Maurya K R, Mudit Dixit

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20167 2025-10-20 cs.CV 50%

Conformal Risk Control for Pulmonary Nodule Detection

Roel Hulsman, Valentin Comte, Lorenzo Bertolini, Tobias Wiesenthal, Antonio Puertas Gallardo, Mario Ceresa

机构 * University of Amsterdam(阿姆斯特丹大学) European Commission, Joint Research Centre (JRC)(欧洲委员会联合研究中心)

专题命中 安全评测 :safety(abstract)

Journal ref Proceedings of the Fourteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 266:445-463, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏