arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2507.19455 2025-07-28 cs.LG 57%

Forest-Guided Clustering -- Shedding Light into the Random Forest Black Box

Lisa Barros de Andrade e Sousa, Gregor Miller, Ronan Le Gleut, Dominik Thalmeier, Helena Pelin, Marie Piraud

机构 * Helmholtz AI(海德堡人工智能研究所) Helmholtz Munich(海德堡慕尼黑)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19174 2025-07-28 cs.LG 57%

Automatic Cough Analysis for Non-Small Cell Lung Cancer Detection

Chiara Giangregorio, Cristina Maria Licciardello, Vanja Miskovic, Leonardo Provenzano, Alessandra Laura Giulia Pedrocchi, Andra Diana Dumitrascu, Arsela Prelaj, Marina Chiara Garassino, Emilia Ambrosini, Simona Ferrante

机构 * Department of Electronics, Information and Bioengineering, Politecnico di Milano(电子、信息与生物工程学院,米兰理工学院) Fondazione IRCCS Istituto Nazionale dei Tumori di Milano(米兰国家肿瘤研究所) Department of Medicine, Section of Hematology/Oncology, University of Chicago(医学学院,血液学/肿瘤学部门,芝加哥大学) LEARNLab, IRCCS Istituto Neurologico Carlo Besta(LEARN实验室,卡尔·贝斯塔神经病学研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Emilia Ambrosini and Simona Ferrante equally contributed to the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18667 2025-07-28 cs.CV cs.AI 57%

Gen-AI Police Sketches with Stable Diffusion

Nicholas Fidalgo, Aaron Contreras, Katherine Harvey, Johnny Ni

机构 * Harvard College(哈佛学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01482 2025-07-28 cs.AI 57%

Understanding LLM Scientific Reasoning through Promptings and Model's Explanation on the Answers

Alice Rueda, Mohammed S. Hassan, Argyrios Perivolaris, Bazen G. Teferra, Reza Samavi, Sirisha Rambhatla, Yuqi Wu, Yanbo Zhang, Bo Cao, Divya Sharma, Sridhar Krishnan, Venkat Bhat

机构 * University of Toronto Department of Psychiatry(多伦多大学精神病学系) Toronto Metropolitan University(多伦多 Metropolitan 大学) St. Michael’s Hospital, Unity Health Toronto(圣米歇尔医院,统一健康多伦多) Department of Electrical, Computer, and Biomedical Engineering(电气、计算机和生物医学工程系) University of Waterloo(滑铁卢大学) Department of Management Science and Engineering(管理科学与工程系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17147 2025-07-24 cs.CL 57%

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards

Cheng Liu, Yifei Lu, Fanghua Ye, Jian Li, Xingyu Chen, Feiliang Ren, Zhaopeng Tu, Xiaolong Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18489 2025-07-24 cs.AI cs.ET cs.RO cs.SE 57%

LLM as a code generator in Agile Model Driven Development

Ahmed R. Sadik, Sebastian Brulin, Markus Olhofer

机构 * Honda Research Institute Europe(本田欧洲研究机构)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16572 2025-07-23 cs.CL 57%

Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models

Mohamad Ballout, Serwan Jassim, Elia Bruni

机构 * Institute of Cognitive Science, University of Osnabrück(认知科学研究所,奥斯纳布吕克大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15761 2025-07-22 cs.AI 57%

GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts

Jingyi Zheng, Zifan Peng, Yule Liu, Junfeng Wang, Yifan Liao, Wenhan Dong, Xinlei He

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15239 2025-07-22 cs.AI eess.SP 57%

Explainable Artificial Intelligence based Soft Evaluation Indicator for Arc Fault Diagnosis

Qianchao Wang, Yuxuan Ding, Chuanzhen Jia, Zhe Li, Yaping Du

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15198 2025-07-22 cs.CL 57%

Collaborative Distillation Strategies for Parameter-Efficient Language Model Deployment

Xiandong Meng, Yan Wu, Yexin Tian, Xin Hu, Tianze Kang, Junliang Du

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07057 2025-07-22 cs.CL 57%

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş, Sercan Karakaş, Banu Diri, Savaş Yıldırım

机构 * Yıldız Technical University(伊兹密尔技术大学) Yeditepe University(耶迪特佩大学) University of Chicago(芝加哥大学) Istanbul Bilgi University(伊斯坦布尔比金大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14807 2025-07-22 cs.CV cs.AI 57%

Seeing Through Deepfakes: A Human-Inspired Framework for Multi-Face Detection

Juan Hu, Shaojing Fan, Terence Sim

机构 * National University of Singapore(新加坡国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14513 2025-07-22 cs.AI 57%

Amico: An Event-Driven Modular Framework for Persistent and Embedded Autonomy

Hongyi Yang, Yue Pan, Jiayi Xu, Kelsen Liu

机构 * Department of Aeronautics and Astronautics, Zhejiang University(浙江大学航空航天学院) Department of Computer Science, University College London(伦敦大学学院计算机系) Steinhardt School of Culture, Education, and Human Development, New York University(纽约大学文化、教育与人类发展学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14492 2025-07-22 cs.LG stat.ML 57%

Glitches in Decision Tree Ensemble Models

Satyankar Chandra, Ashutosh Gupta, Kaushik Mallik, Krishna Shankaranarayanan, Namrita Varshney

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14355 2025-07-22 cs.CL 57%

Can LLMs Infer Personality from Real World Conversations?

Jianfeng Zhu, Ruoming Jin, Karin G. Coifman

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 21 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14107 2025-07-21 cs.AI cs.IR 57%

Automated Interpretation of Non-Destructive Evaluation Contour Maps Using Large Language Models for Bridge Condition Assessment

Viraj Nishesh Darji, Callie C. Liao, Duoduo Liao

机构 * School of Computing(计算学院) George Mason University(乔治·玛莎大学) College of Science(科学学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Journal ref IEEE BigData, Year: 2024; Page: 3258-3263

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13666 2025-07-21 cs.CL 57%

KiC: Keyword-inspired Cascade for Cost-Efficient Text Generation with LLMs

Woo-Chan Kim, Ji-Hoon Park, Seong-Whan Lee

机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17735 2025-07-21 cs.AI 57%

SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator

Xueyang Zhou, Weidong Wang, Lin Lu, Jiawen Shi, Guiyao Tie, Yongtian Xu, Lixing Chen, Pan Zhou, Neil Zhenqiang Gong, Lichao Sun

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 38 pages;12 figures;12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13499 2025-07-21 cs.SE cs.AI cs.PL 57%

AI-Assisted Fixes to Code Review Comments at Scale

Chandra Maddila, Negar Ghorbani, James Saindon, Parth Thakkar, Vijayaraghavan Murali, Rui Abreu, Jingyue Shen, Brian Zhou, Nachiappan Nagappan, Peter C. Rigby

机构 * Meta

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13007 2025-07-18 cs.AI 57%

Exploiting Constraint Reasoning to Build Graphical Explanations for Mixed-Integer Linear Programming

Roger Xavier Lera-Leri, Filippo Bistaffa, Athina Georgara, Juan Antonio Rodriguez-Aguilar

机构 * Artificial Intelligence Research Institute (IIIA-CSIC)(人工智能研究所(IIIA-CSIC)) University of Southampton(南安普顿大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments To appear in Lecture Notes in Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12750 2025-07-18 cs.LG cs.CV 57%

Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning

Suorong Yang, Peijia Li, Yujie Liu, Zhiming Xu, Peng Ye, Wanli Ouyang, Furao Shen, Dongzhan Zhou

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00749 2025-07-18 cs.MA cs.AI 57%

Coral Protocol: Open Infrastructure Connecting The Internet of Agents

Roman J. Georgio, Caelum Forder, Suman Deb, Andri Rahimov, Peter Carroll, Önder Gürcan

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 46 pages, 7 figures, Whitepaper

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12207 2025-07-17 cs.AI cs.NE 57%

BuildEvo: Designing Building Energy Consumption Forecasting Heuristics via LLM-driven Evolution

Subin Lin, Chuanbo Hua

机构 * Department of the Built Environment, National University of Singapore(新加坡国立大学建筑环境系) Department of Industrial and Systems Engineering, Korea Advanced Institute of Science and Technology(韩国科学技术院工业与系统工程系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments ICML 2025 CO-Build Workshop Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09567 2025-07-17 cs.RO cs.AI 57%

Enhancing Trust in Autonomous Agents: An Architecture for Accountability and Explainability through Blockchain and Large Language Models

Laura Fernández-Becerra, Miguel Ángel González-Santamarta, Ángel Manuel Guerrero-Higueras, Francisco Javier Rodríguez-Lera, Vicente Matellán Olivera

机构 * Robotics Group, University of León(莱昂大学机器人组)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00877 2025-07-16 cs.LG 57%

Patch-wise Structural Loss for Time Series Forecasting

Dilfira Kudrat, Zongxia Xie, Yanru Sun, Tianyu Jia, Qinghua Hu

机构 * College of Intelligence and Computing, Tianjin University, China(智能与计算学院,天津大学,中国)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10223 2025-07-15 cs.CV cs.AI 57%

ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis Users

Xiangyu Yin, Boyuan Yang, Weichen Liu, Qiyao Xue, Abrar Alamri, Goeran Fiedler, Wei Gao

机构 * University of Pittsburgh(匹兹堡大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by ICCV'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09701 2025-07-15 cs.CL 57%

MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs

Shulin Huang, Linyi Yang, Yue Zhang

机构 * Zhejiang University(浙江大学) School of Engineering, Westlake University(西溪大学工程学院) Southern University of Science and Technology(南方科技大学) Institute of Advanced Technology, Westlake Institute for Advanced Study(西溪研究院先进技术研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08967 2025-07-15 cs.CL 57%

Self-Improving Model Steering

Rongyi Zhu, Yuhui Wang, Tanqiu Jiang, Jiacheng Liang, Ting Wang

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08869 2025-07-15 cs.CY 57%

Preliminary Analysis of Construction Work Zone on Roadways in Florida by Crash Severity

Tatiana Deslouches, Doreen Kobelo Regalado, Mohamed Khalafalla, Tejal Mulay

专题命中 安全评测 :safety(abstract);分类 cs.CY

Journal ref TRB 104th Annual Meeting, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02948 2025-07-15 cs.CV cs.AI cs.RO 57%

DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction

Zhiyi Hou, Enhui Ma, Fang Li, Zhiyi Lai, Kalok Ho, Zhanqian Wu, Lijun Zhou, Long Chen, Chitian Sun, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Kaicheng Yu

机构 * Westlake University(西湖大学) Xiaomi EV(小米汽车) Zhejiang University(浙江大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 12 pages, 4 figures. Code available at https://github.com/hzy138/DriveMRP

详情

展开后加载摘要…

URL PDF HTML 收藏