arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-29 至 2025-09-29 共收录 15 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 15 篇

2509.10655 2025-09-29 cs.CR cs.CY 83%

Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential

Charankumar Akiri, Harrison Simpson, Kshitiz Aryal, Aarav Khanna, Maanak Gupta

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13577 2025-09-29 cs.SD cs.AI eess.AS 79%

VocalAgent: Large Language Models for Vocal Health Diagnostics with Safety-Aware Evaluation

Yubin Kim, Taehan Kim, Wonjune Kang, Eugene Park, Joonsik Yoon, Dongjae Lee, Xin Liu, Daniel McDuff, Hyeonhoon Lee, Cynthia Breazeal, Hae Won Park

机构 * Massachusetts Institute of Technology, USA(麻省理工学院, 美国) Seoul National University Hospital, South Korea(首尔国立大学医院, 韩国) Doctor Diary, South Korea(Doctor Diary, 韩国) Google Research, USA(谷歌研究, 美国) Independent Researcher(独立研究者)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments Accepted by Proceedings of Interspeech 2025; Website: https://han811.github.io/VocalAgent2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21368 2025-09-29 cs.CV cs.AI 79%

Safety Assessment of Scaffolding on Construction Site using AI

Sameer Prabhu, Amit Patwardhan, Ramin Karim

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22651 2025-09-29 cs.CL cs.AI cs.CV cs.HC cs.SD 73%

VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing

Ke Wang, Houxing Ren, Zimu Lu, Mingjie Zhan, Hongsheng Li

机构 * CUHK MMLab(香港中文大学多媒体实验室) SenseTime Research(商汤科技研究院) CPII under InnoHK(创新香港下的CPII)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05418 2025-09-29 cs.CL cs.AI cs.LG 67%

Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning

Jaedong Hwang, Kumar Tanmay, Seok-Jin Lee, Ayush Agrawal, Hamid Palangi, Kumar Ayush, Ila Fiete, Paul Pu Liang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21579 2025-09-29 cs.LG cs.CL 62%

Leveraging Big Data Frameworks for Spam Detection in Amazon Reviews

Mst Eshita Khatun, Halima Akter, Tasnimul Rehan, Toufiq Ahmed

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

Comments Accepted & presented at THE 16th INTERNATIONAL IEEE CONFERENCE ON COMPUTING, COMMUNICATION AND NETWORKING TECHNOLOGIES (ICCCNT) 2025

Journal ref THE 16th INTERNATIONAL IEEE CONFERENCE ON COMPUTING, COMMUNICATION AND NETWORKING TECHNOLOGIES (ICCCNT) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21117 2025-09-29 cs.AI cs.CL 62%

TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them

Yidong Wang, Yunze Song, Tingyuan Zhu, Xuanwang Zhang, Zhuohao Yu, Hao Chen, Chiyu Song, Qiufeng Wang, Cunxiang Wang, Zhen Wu, Xinyu Dai, Yue Zhang, Wei Ye, Shikun Zhang

机构 * Peking University(北京大学) National University of Singapore(新加坡国立大学) Institute of Science Tokyo(东京科学研究院) Nanjing University(南京大学) Westlake University(西湖大学) Southeast University(东南大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 22 pages, 9 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15253 2025-09-29 cs.CL cs.AI 62%

Conflict-Aware Soft Prompting for Retrieval-Augmented Generation

Eunseong Choi, June Park, Hyeri Lee, Jongwuk Lee

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025; 15 pages; 5 figures, 11 tables; Code available at https://github.com/eunseongc/CARE

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21949 2025-09-29 cs.NI cs.CL 57%

Evaluating Open-Source Large Language Models for Technical Telecom Question Answering

Arina Caraus, Alessio Buscemi, Sumit Kumar, Ion Turcanu

机构 * Luxembourg Institute of Science and Technology (LIST)(卢森堡科学与技术研究院) RMT Labs(RMT实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Accepted at the IEEE GLOBECOM Workshops 2025: "Large AI Model over Future Wireless Networks"

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21749 2025-09-29 cs.CL cs.SD 57%

Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models

Zhen Xiong, Yujun Cai, Zhecheng Li, Junsong Yuan, Yiwei Wang

机构 * University of Southern California(南加州大学) University of Queensland(昆士兰大学) University of California, San Diego(加州大学圣地亚哥分校) University of Buffalo(布法罗大学) University of California, Merced(加州大学默塞德分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21662 2025-09-29 cs.LG 57%

MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning

Afrina Tabassum, Bin Guo, Xiyao Ma, Hoda Eldardiry, Ismini Lourentzou

机构 * Amazon(亚马逊公司) Alexa, Amazon(亚马逊Alexa部门) Virginia Tech(弗吉尼亚理工大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 17 pages, 9 figures, 14 tables, Findings of the Association for Computational Linguistics: EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21450 2025-09-29 cs.CL 57%

LLM-Based Support for Diabetes Diagnosis: Opportunities, Scenarios, and Challenges with GPT-5

Gaurav Kumar Gupta, Nirajan Acharya, Pranal Pande

机构 * Youngstown State University(扬斯敦州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17978 2025-09-29 cs.AI cs.LO 57%

The STAR-XAI Protocol: A Framework for Inducing and Verifying Agency, Reasoning, and Reliability in AI Agents

Antoni Guasch, Maria Isabel Valdez

机构 * Ixent Games

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Version 2: This article consolidates and replaces a previous version to present the complete research in a single, comprehensive manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22228 2025-09-29 cs.CV 50%

UrbanFeel: A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective

Jun He, Yi Lin, Zilong Huang, Jiacong Yin, Junyan Ye, Yuchuan Zhou, Weijia Li, Xiang Zhang

机构 * Sun Yat-sen University(中山大学)

专题命中 安全评测 :safety(abstract)

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03334 2025-09-29 cs.CV cs.DB 50%

OS-W2S: An Automatic Labeling Engine for Language-Guided Open-Set Aerial Object Detection

Guoting Wei, Yu Liu, Xia Yuan, Xizhe Xue, Linlin Guo, Yifan Yang, Chunxia Zhao, Zongwen Bai, Haokui Zhang, Rong Xiao

机构 * Nanjing University of Science and Technology(南京理工大学) Intellifusion Inc.(Intellifusion公司) Northwestern Polytechnical University(西北工业大学) Zhejiang Lab(浙江实验室) Yan’an University(延安大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏