arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-29 至 2025-09-29 共收录 52 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 15 篇

2509.21579 2025-09-29 cs.LG cs.CL 62%

Leveraging Big Data Frameworks for Spam Detection in Amazon Reviews

Mst Eshita Khatun, Halima Akter, Tasnimul Rehan, Toufiq Ahmed

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

Comments Accepted & presented at THE 16th INTERNATIONAL IEEE CONFERENCE ON COMPUTING, COMMUNICATION AND NETWORKING TECHNOLOGIES (ICCCNT) 2025

Journal ref THE 16th INTERNATIONAL IEEE CONFERENCE ON COMPUTING, COMMUNICATION AND NETWORKING TECHNOLOGIES (ICCCNT) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21117 2025-09-29 cs.AI cs.CL 62%

TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them

Yidong Wang, Yunze Song, Tingyuan Zhu, Xuanwang Zhang, Zhuohao Yu, Hao Chen, Chiyu Song, Qiufeng Wang, Cunxiang Wang, Zhen Wu, Xinyu Dai, Yue Zhang, Wei Ye, Shikun Zhang

机构 * Peking University(北京大学) National University of Singapore(新加坡国立大学) Institute of Science Tokyo(东京科学研究院) Nanjing University(南京大学) Westlake University(西湖大学) Southeast University(东南大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 22 pages, 9 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15253 2025-09-29 cs.CL cs.AI 62%

Conflict-Aware Soft Prompting for Retrieval-Augmented Generation

Eunseong Choi, June Park, Hyeri Lee, Jongwuk Lee

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025; 15 pages; 5 figures, 11 tables; Code available at https://github.com/eunseongc/CARE

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21949 2025-09-29 cs.NI cs.CL 57%

Evaluating Open-Source Large Language Models for Technical Telecom Question Answering

Arina Caraus, Alessio Buscemi, Sumit Kumar, Ion Turcanu

机构 * Luxembourg Institute of Science and Technology (LIST)(卢森堡科学与技术研究院) RMT Labs(RMT实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Accepted at the IEEE GLOBECOM Workshops 2025: "Large AI Model over Future Wireless Networks"

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21749 2025-09-29 cs.CL cs.SD 57%

Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models

Zhen Xiong, Yujun Cai, Zhecheng Li, Junsong Yuan, Yiwei Wang

机构 * University of Southern California(南加州大学) University of Queensland(昆士兰大学) University of California, San Diego(加州大学圣地亚哥分校) University of Buffalo(布法罗大学) University of California, Merced(加州大学默塞德分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21662 2025-09-29 cs.LG 57%

MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning

Afrina Tabassum, Bin Guo, Xiyao Ma, Hoda Eldardiry, Ismini Lourentzou

机构 * Amazon(亚马逊公司) Alexa, Amazon(亚马逊Alexa部门) Virginia Tech(弗吉尼亚理工大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 17 pages, 9 figures, 14 tables, Findings of the Association for Computational Linguistics: EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21450 2025-09-29 cs.CL 57%

LLM-Based Support for Diabetes Diagnosis: Opportunities, Scenarios, and Challenges with GPT-5

Gaurav Kumar Gupta, Nirajan Acharya, Pranal Pande

机构 * Youngstown State University(扬斯敦州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17978 2025-09-29 cs.AI cs.LO 57%

The STAR-XAI Protocol: A Framework for Inducing and Verifying Agency, Reasoning, and Reliability in AI Agents

Antoni Guasch, Maria Isabel Valdez

机构 * Ixent Games

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Version 2: This article consolidates and replaces a previous version to present the complete research in a single, comprehensive manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22228 2025-09-29 cs.CV 50%

UrbanFeel: A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective

Jun He, Yi Lin, Zilong Huang, Jiacong Yin, Junyan Ye, Yuchuan Zhou, Weijia Li, Xiang Zhang

机构 * Sun Yat-sen University(中山大学)

专题命中 安全评测 :safety(abstract)

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03334 2025-09-29 cs.CV cs.DB 50%

OS-W2S: An Automatic Labeling Engine for Language-Guided Open-Set Aerial Object Detection

Guoting Wei, Yu Liu, Xia Yuan, Xizhe Xue, Linlin Guo, Yifan Yang, Chunxia Zhao, Zongwen Bai, Haokui Zhang, Rong Xiao

机构 * Nanjing University of Science and Technology(南京理工大学) Intellifusion Inc.(Intellifusion公司) Northwestern Polytechnical University(西北工业大学) Zhejiang Lab(浙江实验室) Yan’an University(延安大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 5 篇

2505.10426 2025-09-29 cs.CY cs.AI cs.HC math.HO 71%

Formalising Human-in-the-Loop: Computational Reductions, Failure Modes, and Legal-Moral Responsibility

Maurice Chiodo, Dennis Müller, Paul Siewert, Jean-Luc Wetherall, Zoya Yasmine, John Burden

机构 * Centre for the Study of Existential Risk(存在风险研究中心) University of Cambridge(剑桥大学) Institute of Mathematics Education(数学教育研究所) University of Cologne(科隆大学) Department of Computer Science and Technology(计算机科学与技术系) DeepFin Research(DeepFin研究) Faculty of Law(法学院) University of Oxford(牛津大学) Leverhulme Centre for the Future of Intelligence(未来智能中心)

专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.AI、cs.CY;trustworthy(comments);AI safety(comments)

Comments 31 pages. Keywords: Human-in-the-loop, Automated decision making system, Human oversight in sociotechnical systems, Oracle machine, AI safety, Trustworthy AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16355 2025-09-29 cs.LG cs.AI 62%

How Strategic Agents Respond: Comparing Analytical Models with LLM-Generated Responses in Strategic Classification

Tian Xie, Pavan Rauch, Xueru Zhang

机构 * The Ohio State University(俄亥俄州立大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Add GPT 5 experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17805 2025-09-29 cs.CY cs.AI 62%

Biospheric AI

Marcin Korecki

机构 * TU Delft(代尔夫特理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21542 2025-09-29 cs.HC cs.AI 57%

Psychological and behavioural responses in human-agent vs. human-human interactions: a systematic review and meta-analysis

Jianan Zhou, Fleur Corbett, Joori Byun, Talya Porat, Nejra van Zalk

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22424 2025-09-29 q-bio.OT 50%

Desiderata for a biomedical knowledge network: opportunities, challenges and future Directions

Chunlei Wu, Hongfang Liu, Jason Flannick, Mark A. Musen, Andrew I. Su, Lawrence Hunter, Thomas M. Powers, Cathy H. Wu

专题命中 AI治理与伦理 :trustworthy(abstract)

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 7 篇

2505.17030 2025-09-29 eess.IV cs.LG 79%

Distillation-Enabled Knowledge Alignment Protocol for Semantic Communication in AI Agent Networks

Jingzhi Hu, Geoffrey Ye Li

机构 * Department of Electrical and Electronic Engineering, Imperial College London(帝国理工学院伦敦分校电子与电气工程系)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Code available at https://github.com/DJ-Duke/DeKAP

Journal ref IEEE Communications Letters, early access, Aug. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21665 2025-09-29 cs.HC 78%

Alignment Without Understanding: A Message- and Conversation-Centered Approach to Understanding AI Sycophancy

Lihua Du, Xing Lyu, Lezi Xie, Bo Feng

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21358 2025-09-29 cs.CV cs.AI 74%

MDF-MLLM: Deep Fusion Through Cross-Modal Feature Alignment for Contextually Aware Fundoscopic Image Classification

Jason Jordan, Mohammadreza Akbari Lor, Peter Koulen, Mei-Ling Shyu, Shu-Ching Chen

专题命中 其他安全 :alignment(title);分类 cs.AI

Comments Word count: 5157, Table count: 2, Figure count: 5

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21625 2025-09-29 cs.SD cs.AI cs.LG eess.AS 62%

Guiding Audio Editing with Audio Language Model

Zitong Lan, Yiduo Hao, Mingmin Zhao

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22230 2025-09-29 cs.CL 57%

In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners

Jaehoon Kim, Kwangwook Seo, Dongha Lee

机构 * Yonsei University(延世大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15215 2025-09-29 cs.LG 57%

Multi-Channel Differential Transformer for Cross-Domain Sleep Stage Classification with Heterogeneous EEG and EOG

Benjamin Wei Hao Chin, Yuin Torng Yew, Haocheng Wu, Lanxin Liang, Chow Khuen Chan, Norita Mohd Zain, Siti Balqis Samdin, Sim Kuan Goh

机构 * School of Artificial Intelligence and Robotics, Xiamen University Malaysia(厦门大学马来西亚分校人工智能与机器人学院) Department of Biomedical Engineering, Universiti Malaya(马来亚大学生物医学工程系) Department of Biomedical Engineering and Health Sciences, Universiti Teknologi Malaysia(马来西亚理工学院生物医学工程与健康科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments SleepDIFFormer 8 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21556 2025-09-29 cs.CL 57%

VAT-KG: Knowledge-Intensive Multimodal Knowledge Graph Dataset for Retrieval-Augmented Generation

Hyeongcheol Park, Jiyoung Seo, MinHyuk Jang, Hogun Park, Ha Dam Baek, Gyusam Chang, Hyeonsoo Im, Sangpil Kim

机构 * Korea University(韩国大学) Sungkyunkwan University(全北大学) Hanwha Systems(韩华系统)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Project Page: https://vatkg.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏