arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-13 至 2025-08-13 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 10 篇

2501.13983 2025-08-13 cs.CL cs.AI 81%

AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models

Yang Fan

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments There are serious academic problems in this paper, such as data falsification and plagiarism in the method of the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02078 2025-08-13 cs.CV cs.AI cs.LG 81%

From Lab to Field: Real-World Evaluation of an AI-Driven Smart Video Solution to Enhance Community Safety

Shanle Yao, Babak Rahimi Ardabili, Armin Danesh Pazho, Ghazal Alinezhad Noghre, Christopher Neff, Lauren Bourque, Hamed Tabkhi

机构 * Department of Electrical and Computer Engineering, University of North Carolina at Charlotte(电气与计算机工程系,北卡罗来纳大学夏洛特分校) Department of Public Policy, University of North Carolina at Charlotte(公共政策系,北卡罗来纳大学夏洛特分校)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08273 2025-08-13 cs.CL cs.LG 81%

TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning

Kristian Miok, Blaz Škrlj, Daniela Zaharie, Marko Robnik Šikonja

机构 * Faculty of Computer and Information Science, University of Ljubljana, Slovenia(卢布尔雅那大学计算机与信息科学学院) ICAM - Advanced Environmental Research Institute, West University of Timisoara, Romania(蒂米șoara西大学先进环境研究所) Department of Computer Science, West University of Timisoara, Romania(蒂米șoa拉西大学计算机科学系)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08777 2025-08-13 cs.IR cs.AI cs.LG 62%

Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge

Francesco Fabbri, Gustavo Penha, Edoardo D'Amico, Alice Wang, Marco De Nadai, Jackie Doremus, Paul Gigioli, Andreas Damianou, Oskar Stal, Mounia Lalmas

机构 * Spotify

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted at RecSys '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08629 2025-08-13 cs.CY cs.AI 62%

Securing Educational LLMs: A Generalised Taxonomy of Attacks on LLMs and DREAD Risk Assessment

Farzana Zahid, Anjalika Sewwandi, Lee Brandon, Vimal Kumar, Roopak Sinha

机构 * University of Waikato(怀卡托大学) Deakin University(迪金大学)

专题命中 安全评测 :jailbreak(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08277 2025-08-13 cs.CL cs.LG 62%

Objective Metrics for Evaluating Large Language Models Using External Data Sources

Haoze Du, Richard Li, Edward Gehringer

机构 * Department of Computer Science(计算机科学系) North Carolina State University(北卡罗来纳州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments This version of the paper is lightly revised from the EDM 2025 proceedings for the sake of clarity

Journal ref EDM 2025 Palermo, Italy, July, 2025, pp. 489-495. International Educational Data Mining Society (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06555 2025-08-13 cs.CV cs.CY cs.MA 57%

StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback

Hongbo Ma, Fei Shen, Hongbin Xu, Xiaoce Wang, Gang Xu, Jinkai Zheng, Liangqiong Qu, Ming Li

专题命中 安全评测 :alignment(abstract);分类 cs.CY

Comments 24pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10066 2025-08-13 cs.MM cs.CV 50%

LayLens: Improving Deepfake Understanding through Simplified Explanations

Abhijeet Narang, Parul Gupta, Liuyijia Su, Abhinav Dhall

机构 * Monash University(墨尔本大学)

专题命中 安全评测 :trustworthy(abstract)

Comments Accepted to ACM ICMI 2025 Demos

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19316 2025-08-13 cs.MA 50%

Making Teams and Influencing Agents: Efficiently Coordinating Decision Trees for Interpretable Multi-Agent Reinforcement Learning

Rex Chen, Stephanie Milani, Zhicheng Zhang, Norman Sadeh, Fei Fang

专题命中 安全评测 :safety(abstract)

Comments 17 pages; 2 tables; 12 figures; accepted version, published at the 8th AAAI/ACM Conference on AI, Ethics and Society (AIES '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11002 2025-08-13 cs.SD cs.MM eess.AS 50%

Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation

Yan Rong, Shan Yang, Chenxing Li, Dong Yu, Li Liu

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏