arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-05 至 2025-08-05 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 11 篇

2507.22940 2025-08-05 cs.CL cs.AI 76%

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes

Rui Jiao, Yue Zhang, Jinku Li

机构 * School of Cyber Engineering, Xidian University(西安电子科技大学电子工程学院) School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院)

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01198 2025-08-05 cs.CL cs.AI 62%

Adaptive Content Restriction for Large Language Models via Suffix Optimization

Yige Li, Peihai Jiang, Jun Sun, Peng Shu, Tianming Liu, Zhen Xiang

机构 * Singapore Management University(新加坡管理大学) The University of Georgia(佐治亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02525 2025-08-05 cs.AI 57%

Accurate and Interpretable Postmenstrual Age Prediction via Multimodal Large Language Model

Qifan Chen, Jin Cui, Cindy Duan, Yushuo Han, Yifei Shi

机构 * King’s College London(伦敦国王学院) Imperial College London(帝国理工学院) Columbia University(哥伦比亚大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Submitted to the NeurIPS 2025 Workshop GenAI4Health. Conference website: https://aihealth.ischool.utexas.edu/GenAI4HealthNeurips2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01370 2025-08-05 cs.CL cs.IR 57%

MaRGen: Multi-Agent LLM Approach for Self-Directed Market Research and Analysis

Roman Koshkin, Pengyu Dai, Nozomi Fujikawa, Masahito Togami, Marco Visentini-Scarzanella

机构 * Okinawa Institute of Science and Technology(冲绳科学技术研究所) Institute of Science Tokyo(东京科学研究所) Amazon(亚马逊)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12441 2025-08-05 cs.CV cs.LG 57%

Describe Anything Model for Visual Question Answering on Text-rich Images

Yen-Linh Vu, Dinh-Thang Duong, Truong-Binh Duong, Anh-Khoi Nguyen, Thanh-Huy Nguyen, Le Thien Phuc Nguyen, Jianhua Xing, Xingjian Li, Tianyang Wang, Ulas Bagci, Min Xu

机构 * AI VIETNAM Lab(AI越南实验室) Carnegie Mellon University(卡内基梅隆大学) University of Wisconsin - Madison(威斯康星大学麦迪逊分校) University of Pittsburgh(匹兹堡大学) University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) Northwestern University(西北大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 11 pages, 5 figures. Accepted to VisionDocs @ ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10130 2025-08-05 cs.AI 57%

A Conjecture on a Fundamental Trade-Off between Certainty and Scope in Symbolic and Generative AI

Luciano Floridi

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments version 3

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06762 2025-08-05 cs.LG cs.RO 57%

Investigating Robotaxi Crash Severity with Geographical Random Forest and the Urban Environment

Junfeng Jiao, Seung Gyu Baik, Seung Jun Choi, Yiming Xu

机构 * Urban Information Lab, The University of Texas at Austin(德克萨斯大学奥斯汀分校城市信息实验室)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10047 2025-08-05 cs.MA 50%

Bearing-Distance Flocking with Zone-Based Interactions in Constrained Dynamic Environments

Hossein B. Jond

专题命中 安全评测 :alignment(abstract)

Comments Video for Figure 4: https://youtu.be/5nSml5F2oQk?si=X0q51SXcZiTRBdFS Video for Figure 6 (2D): https://youtu.be/ANkAZ9FMq0o?si=px8oSAeKkUwBR7uf Video for Figure 6 (3D): https://youtu.be/AUlcCH7P73U?si=QFYYeCGLIGdsJCad

Journal ref Journal of Computational Science, Volume 87, 2025, 102574

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02082 2025-08-05 cs.CV 50%

S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework

Yingshu Li, Yunyi Liu, Zhanyu Wang, Xinyu Liang, Lingqiao Liu, Lei Wang, Luping Zhou

机构 * University of Sydney(悉尼大学) University of Wollongong(沃林戈大学) University of Adelaide(阿德莱德大学) First Clinical Medical College, Guangzhou University of Chinese Medicine(广州中医药大学第一临床学院)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01582 2025-08-05 cs.CV 50%

Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models

Xinhui Li, Xinyu He, Qiming Hu, Xiaojie Guo

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18645 2025-08-05 cs.CV 50%

Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings

Azim Ospanov, Mohammad Jalali, Farzan Farnia

机构 * The Chinese University of Hong Kong, Department of Computer Science & Engineering(香港中文大学计算机科学与工程系)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏