arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-08 至 2025-10-08 共收录 16 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 16 篇

2503.10663 2025-10-08 q-bio.NC cs.AI cs.CV cs.LG 81%

Optimal Transport for Brain-Image Alignment: Unveiling Redundancy and Synergy in Neural Information Processing

Yang Xiao, Wang Lu, Jie Ji, Ruimeng Ye, Gen Li, Xiaolong Ma, Bo Hui

机构 * University of Tulsa(图拉大学) Tsinghua University(清华大学) Clemson University(克莱姆森大学) The University of Arizona(亚利桑那大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 14pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06093 2025-10-08 cs.AI 79%

Classical AI vs. LLMs for Decision-Maker Alignment in Health Insurance Choices

Mallika Mainali, Harsha Sureshbabu, Anik Sen, Christopher B. Rauch, Noah D. Reifsnyder, John Meyer, J. T. Turner, Michael W. Floyd, Matthew Molineaux, Rosina O. Weber

机构 * Information Science, Drexel University, Philadelphia, PA 19104 USA Parallax Advanced Research, 4035 Colonel Glenn Hwy, Beavercreek, OH 45431 USA Knexus Research, 174 Waterfront Street, Suite 310, National Harbor, Oxon Hill, MD 20745 USA Information Science \& Computer Science, Drexel University, Philadelphia, PA 19104 USA

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments 15 pages, 3 figures. Accepted at the Twelfth Annual Conference on Advances in Cognitive Systems (ACS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23058 2025-10-08 cs.AI cs.LG 73%

Risk Profiling and Modulation for LLMs

Yikai Wang, Xiaocheng Li, Guanting Chen

机构 * Department of Statistics and Operations Research, UNC-Chapel Hill(统计与运筹学系,北卡罗来纳大学 Chapel Hill 分校) Imperial College Business School, Imperial College London(帝国理工学院伦敦校区商学院)

专题命中 安全评测 :alignment(abstract);RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00192 2025-10-08 cs.CV 71%

Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety

Younggun Kim, Sirnam Swetha, Fazil Kagdi, Mubarak Shah

机构 * Center For Research in Computer Vision, University of Central Florida, USA(计算机视觉研究中心,中央佛罗里达大学) Department of Civil Environmental and Construction Engineering, University of Central Florida, USA(土木环境与建设工程系,中央佛罗里达大学) Department of Computer Science, University of Central Florida, USA(计算机科学系,中央佛罗里达大学)

专题命中 安全评测 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05976 2025-10-08 cs.CV cs.AI cs.LG 62%

Diffusion Models for Low-Light Image Enhancement: A Multi-Perspective Taxonomy and Performance Analysis

Eashan Adhikarla, Yixin Liu, Brian D. Davison

机构 * Lehigh University(莱维大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05972 2025-10-08 cs.CL cs.AI 62%

LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language

Periklis Mantenoglou, Rishi Hazra, Pedro Zuidberg Dos Martires, Luc De Raedt

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05310 2025-10-08 cs.CL cs.AI 62%

RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts

Yining She, Daniel W. Peterson, Marianne Menglin Liu, Vikas Upadhyay, Mohammad Hossein Chaghazardi, Eunsuk Kang, Dan Roth

机构 * Carnegie Mellon University(卡内基梅隆大学) Oracle Cloud Infrastructure(Oracle 云基础设施) University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12543 2025-10-08 cs.AI cs.CV cs.LG 62%

Human + AI for Accelerating Ad Localization Evaluation

Harshit Rajgarhia, Shivali Dalmia, Mengyang Zhao, Mukherji Abhishek, Kiran Ganesh

机构 * Centific Global Solutions Inc.(Centific全球解决方案公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11676 2025-10-08 cs.LG cs.AI stat.ME stat.ML 62%

SKADA-Bench: Benchmarking Unsupervised Domain Adaptation Methods with Realistic Validation On Diverse Modalities

Yanis Lalou, Théo Gnassounou, Antoine Collas, Antoine de Mathelin, Oleksii Kachaiev, Ambroise Odonnat, Alexandre Gramfort, Thomas Moreau, Rémi Flamary

机构 * École Polytechnique, IP Paris, CMAP, UMR 7641(巴黎理工学院) Université Paris-Saclay, Inria, CEA(巴黎萨克雷大学) Inria(法国国家信息与自动化技术研究院) CEA(法国原子能机构) Centre Borelli, ENS Paris-Saclay(巴黎-萨克雷大学博雷利中心) Università degli Studi di Genova(热那亚大学) Inria, Univ. Rennes 2, CNRS, IRISA(法国国家信息与自动化技术研究院、里昂二大学、法国国家科学研究中心、IRISA)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Published in Transactions on Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03413 2025-10-08 cs.CE cs.AI 57%

Report of the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science

Lois Curfman McInnes, Dorian Arnold, Prasanna Balaprakash, Mike Bernhardt, Beth Cerny, Anshu Dubey, Roscoe Giles, Denice Ward Hood, Mary Ann Leung, Vanessa Lopez-Marrero, Paul Messina, Olivia B. Newton, Chris Oehmen, Stefan M. Wild, Jim Willenbring, Lou Woodley, Tony Baylis, David E. Bernholdt, Chris Camano, Johannah Cohoon, Charles Ferenbaugh, Stephen M. Fiore, Sandra Gesing, Diego Gomez-Zara, James Howison, Tanzima Islam, David Kepczynski, Charles Lively, Harshitha Menon, Bronson Messer, Marieme Ngom, Umesh Paliath, Michael E. Papka, Irene Qualters, Elaine M. Raybourn, Katherine Riley, Paulina Rodriguez, Damian Rouson, Michelle Schwalbe, Sudip K. Seal, Ozge Surer, Valerie Taylor, Lingfei Wu

机构 * Argonne National Laboratory(阿贡国家实验室) Emory University(埃默里大学) Oak Ridge National Laboratory(橡树岭国家实验室) Team Libra(团队Libra) Boston University(波士顿大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Sustainable Horizons Institute(可持续远景研究所) Stony Brook University(石溪大学) University of Montana(蒙大拿大学) Pacific Northwest National Laboratory(太平洋西北国家实验室) Lawrence Berkeley National Laboratory(伯克利国家实验室) Sandia National Laboratories(桑塔那国家实验室) Center for Scientific Collaboration and Community Engagement(科学协作与社区参与中心) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) Californi(加利福尼亚)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 38 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05519 2025-10-08 cs.CY 57%

Assessing Human Rights Risks in AI: A Framework for Model Evaluation

Vyoma Raman, Camille Chabot, Betsy Popken

专题命中 安全评测 :safety(abstract);分类 cs.CY

Comments AAAI/ACM Conference on AI, Ethics, and Society (AIES) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05185 2025-10-08 cs.MA cs.CE cs.CY cs.NE cs.SI 57%

AgentZero++: Modeling Fear-Based Behavior

Vrinda Malhotra, Jiaman Li, Nandini Pisupati

专题命中 安全评测 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05142 2025-10-08 cs.CL cond-mat.mtrl-sci 57%

Reliable End-to-End Material Information Extraction from the Literature with Source-Tracked Multi-Stage Large Language Models

Xin Wang, Anshu Raj, Matthew Luebbe, Haiming Wen, Shuozhi Xu, Kun Lu

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 27 pages, 4 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04093 2025-10-08 cs.AI 57%

Harnessing LLM for Noise-Robust Cognitive Diagnosis in Web-Based Intelligent Education Systems

Guixian Zhang, Guan Yuan, Ziqi Xu, Yanmei Zhang, Jing Ren, Zhenyun Deng, Debo Cheng

机构 * School of Computer Science and Technology/School of Artificial Intelligence, China University of Mining and Technology(计算机科学与技术学院/人工智能学院,中国矿业大学) School of Computing Technologies, RMIT University(计算技术学院,拉筹伯大学) Department of Computer Science and Technology, University of Cambridge(计算机科学与技术系,剑桥大学) School of Computer Science and Technology, Hainan University(计算机科学与技术学院,海南大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18427 2025-10-08 q-fin.CP q-fin.RM 50%

Tracing Positional Bias in Financial Decision-Making: Mechanistic Insights from Qwen2.5

Fabrizio Dimino, Krati Saxena, Bhaskarjit Sarmah, Stefano Pasquali

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00815 2025-10-08 econ.TH cs.GT math.OC 50%

Measurement of Trustworthiness of the Online Reviews

Dipankar Das

专题命中 安全评测 :trustworthy(abstract)

Comments This is a minor revision version and considers some intuitions related to applications. Moreover, a detailed algorithm has been added to facilitate a better understanding

详情

展开后加载摘要…

URL PDF HTML 收藏