arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9451 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9451 篇

2503.07376 2025-10-23 eess.SY cs.RO cs.SY 50%

AttentionSwarm: Reinforcement Learning with Attention Control Barier Function for Crazyflie Drones in Dynamic Environments

Grik Tadevosyan, Valerii Serpiva, Aleksey Fedoseev, Roohan Ahmed Khan, Demetros Aschu, Faryal Batool, Nickolay Efanov, Artem Mikhaylov, Dzmitry Tsetserukou

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Intelligent Space Robotics Laboratory(智能空间机器人实验室) Skolkovo Institute of Science and Technology(斯克尔科维科学与技术学院) AI Center(人工智能中心) Moscow Institute of Physics and Technology(莫斯科物理技术学院)

专题命中 安全评测 :safety(abstract)

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09814 2025-10-23 cs.GR cs.CV cs.SD eess.AS 50%

Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis

Zeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao, Chuan Lin, Baoquan Chen, Libin Liu

机构 * School of Electronics Engineering and Computer Science, Peking University(电子工程与计算机科学学院,北京大学) School of Computer Science, Peking University(计算机学院,北京大学) Renmin University of China(中国人民大学) Shandong University(山东大学) Peking University(北京大学) State Key Lab of General AI(通用人工智能国家重点实验室)

专题命中 安全评测 :alignment(abstract)

Comments SIGGRAPH 2024 (Journal Track); Project page: https://pku-mocca.github.io/Semantic-Gesticulator-Page

Journal ref ACM Transactions on Graphics (TOG) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18563 2025-10-22 cs.CR 50%

The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability

Zijie Xu, Minfeng Qi, Shiqing Wu, Lefeng Zhang, Qiwen Wei, Han He, Ningran Li

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18364 2025-10-22 cs.IR cs.SE 50%

Evaluating LLM-Based Mobile App Recommendations: An Empirical Study

Quim Motger, Xavier Franch, Vincenzo Gervasi, Jordi Marco

专题命中 安全评测 :alignment(abstract)

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18169 2025-10-22 eess.AS cs.SD 50%

Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction

Yu-Wen Chen, William Ho, Sasha M. Vergez, Grace Flaherty, Pallavi Gupta, Zhihong Zhang, Maryam Zolnoori, Margaret V. McDonald, Maxim Topaz, Zoran Kostic, Julia Hirschberg

机构 * The Fu Foundation School of Engineering and Applied Science, Columbia University(哥伦比亚大学福基金会工程与应用科学学院) School of Nursing, Columbia University(哥伦比亚大学护理学院) Center for Home Care Policy & Research, VNS Health(VNS健康居家护理政策与研究中心)

专题命中 安全评测 :alignment(abstract)

Comments The Second Workshop on GenAI for Health at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02236 2025-10-22 cs.CV cs.MM cs.SD eess.AS 50%

3D Audio-Visual Segmentation

Artem Sokolov, Swapnil Bhosale, Xiatian Zhu

机构 * University of Surrey, UK(Surrey大学)

专题命中 安全评测 :alignment(abstract)

Comments Accepted at the NeurIPS 2024 Workshop on Audio Imagination; this version updates the project page link

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17034 2025-10-21 cs.CV 50%

Where, Not What: Compelling Video LLMs to Learn Geometric Causality for 3D-Grounding

Yutong Zhong

机构 * New York University(纽约大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20149 2025-10-21 cs.SE 50%

Synergistic Enhancement of Requirement-to-Code Traceability: A Framework Combining Large Language Model based Data Augmentation and an Advanced Encoder

Jianzhang Zhang, Jialong Zhou, Nan Niu, Jinping Hua, Chuang Liu

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01879 2025-10-21 cs.MM cs.CV cs.SD eess.AS 50%

Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision

Che Liu, Yingji Zhang, Dong Zhang, Weijie Zhang, Chenggong Gong, Yu Lu, Shilin Zhou, Ziliang Gan, Ziao Wang, Haipang Wu, Ji Liu, André Freitas, Qifan Wang, Zenglin Xu, Rongjuncheng Zhang, Yong Dai

机构 * Imperial College London(伦敦帝国学院) University of Manchester(曼彻斯特大学) HiThink Research(HiThink研究院) Soochow University(苏州大学) Hong Kong Baptist University(香港 Baptist大学) Idiap Research Institute(Idiap研究 institute) Meta AI Fudan University(复旦大学)

专题命中 安全评测 :alignment(abstract)

Comments Project: https://github.com/HiThink-Research/NEXUS-O

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16597 2025-10-21 cs.IR 50%

FRONTIER-RevRec: A Large-scale Dataset for Reviewer Recommendation

Qiyao Peng, Chen Wang, Yinghui Wang, Hongtao Liu, Xuan Guo, Wenjun Wang

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16579 2025-10-21 cs.SE 50%

Human-Aligned Code Readability Assessment with Large Language Models

Wendkûuni C. Ouédraogo, Yinghua Li, Xueqi Dang, Pawel Borsukiewicz, Xin Zhou, Anil Koyuncu, Jacques Klein, David Lo, Tegawendé F. Bissyandé

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06767 2025-10-21 cs.SE 50%

Beyond Surface Similarity: Evaluating LLM-Based Test Refactorings with Structural and Semantic Awareness

Wendkûuni C. Ouédraogo, Yinghua Li, Xueqi Dang, Xin Zhou, Anil Koyuncu, Jacques Klein, David Lo, Tegawendé F. Bissyandé

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15729 2025-10-20 cs.IR 50%

FACE: A General Framework for Mapping Collaborative Filtering Embeddings into LLM Tokens

Chao Wang, Yixin Song, Jinhui Ye, Chuan Qin, Dazhong Shen, Lingfeng Liu, Xiang Wang, Yanyong Zhang

专题命中 安全评测 :alignment(abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15564 2025-10-20 cs.CV 50%

Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation

Xiaoming Zhu, Xu Huang, Qinghongbing Xie, Zhi Deng, Junsheng Yu, Yirui Guan, Zhongyuan Liu, Lin Zhu, Qijun Zhao, Ligang Liu, Long Zeng

机构 * Tsinghua University(清华大学) Tencent(腾讯) Southeast University(东南大学) University of Science and Technology of China(中国科学技术大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15397 2025-10-20 cond-mat.mtrl-sci 50%

Unravelling the Catalytic Activity of Dual-Metal Doped N6-Graphene for Sulfur Reduction via Machine Learning-Accelerated First-Principles Calculations

Sahil Kumar, Adithya Maurya K R, Mudit Dixit

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20167 2025-10-20 cs.CV 50%

Conformal Risk Control for Pulmonary Nodule Detection

Roel Hulsman, Valentin Comte, Lorenzo Bertolini, Tobias Wiesenthal, Antonio Puertas Gallardo, Mario Ceresa

机构 * University of Amsterdam(阿姆斯特丹大学) European Commission, Joint Research Centre (JRC)(欧洲委员会联合研究中心)

专题命中 安全评测 :safety(abstract)

Journal ref Proceedings of the Fourteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 266:445-463, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12995 2025-10-16 cs.CV 50%

Brought a Gun to a Knife Fight: Modern VFM Baselines Outgun Specialized Detectors on In-the-Wild AI Image Detection

Yue Zhou, Xinan He, Kaiqing Lin, Bing Fan, Feng Ding, Jinhua Zeng, Bin Li

机构 * Guangdong Provincial Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室) Shenzhen Key Laboratory of Media Security(深圳媒体安全重点实验室) SZU AFS Joint Innovation Center for AI Technology, Shenzhen University(深圳大学AFS人工智能技术联合创新中心) University of North Texas(北卡罗来纳州立大学) Academy of Forensic Science(法医科学研究院) Nanchang University(南昌大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16902 2025-10-16 cs.CV cs.RO 50%

RealEngine: Simulating Autonomous Driving in Realistic Context

Junzhe Jiang, Nan Song, Jingyu Li, Xiatian Zhu, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院) University of Surrey(萨里大学)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12190 2025-10-15 cs.CV 50%

Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos

Shingo Yokoi, Kento Sasaki, Yu Yamaguchi

机构 * Turing Inc.(图灵公司)

专题命中 安全评测 :safety(abstract)

Comments 2nd Place Winner, ICCV 2025 2COOOL Competition

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11534 2025-10-14 cs.RO cs.SY eess.SY 50%

IntersectioNDE: Learning Complex Urban Traffic Dynamics based on Interaction Decoupling Strategy

Enli Lin, Ziyuan Yang, Qiujing Lu, Jianming Hu, Shuo Feng

机构 * Department of Mechanical Engineering, Tsinghua University(清华大学机械工程系) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 安全评测 :safety(abstract)

Comments Accepted by ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08789 2025-10-14 cs.CV 50%

Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization

Shuo Xing, Soumik Dey, Mingyang Wu, Ashirbad Mishra, Naveen Ravipati, Binbin Li, Hansi Wu, Zhengzhong Tu

机构 * Department of Computer Science(计算机科学系) Cranberry-Lemon University(Cranberry-Lemon 大学) Texas A&M University(德克萨斯A&M大学) eBay Inc.(eBay公司)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11414 2025-10-14 cs.CR 50%

Uncertainty-Aware, Risk-Adaptive Access Control for Agentic Systems using an LLM-Judged TBAC Model

Charles Fleming, Ashish Kundu, Ramana Kompella

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10037 2025-10-14 cs.CE 50%

Automated Glaucoma Report Generation via Dual-Attention Semantic Parallel-LSTM and Multimodal Clinical Data Integration

Cheng Huang, Weizheng Xie, Zeyu Han, Tsengdar Lee, Karanjit Kooner, Jui-Ka Wang, Ning Zhang, Jia Zhang

专题命中 安全评测 :alignment(abstract)

Comments Accepted by IEEE 25th BIBE

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09320 2025-10-13 cs.CV 50%

Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation

Wenyao Zhang, Hongsi Liu, Bohan Li, Jiawei He, Zekun Qi, Yunnan Wang, Shengyang Zhao, Xinqiang Yu, Wenjun Zeng, Xin Jin

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能 MOE 实验室,上海交通大学人工智能学院) Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo, China(宁波数字孪生研究所,东技术研究所,宁波,中国) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Ningbo, China(宁波空间智能与数字衍生关键实验室,宁波,中国) University of Science and Technology of China(中国科学技术大学) CASIA Tsinghua University(清华大学)

专题命中 安全评测 :alignment(abstract)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08003 2025-10-10 cs.CV 50%

CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning

Weihuang Lin, Yiwei Ma, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学)

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18427 2025-10-08 q-fin.CP q-fin.RM 50%

Tracing Positional Bias in Financial Decision-Making: Mechanistic Insights from Qwen2.5

Fabrizio Dimino, Krati Saxena, Bhaskarjit Sarmah, Stefano Pasquali

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00815 2025-10-08 econ.TH cs.GT math.OC 50%

Measurement of Trustworthiness of the Online Reviews

Dipankar Das

专题命中 安全评测 :trustworthy(abstract)

Comments This is a minor revision version and considers some intuitions related to applications. Moreover, a detailed algorithm has been added to facilitate a better understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03909 2025-10-07 cs.CV 50%

Generating Human Motion Videos using a Cascaded Text-to-Video Framework

Hyelin Nam, Hyojun Go, Byeongjun Park, Byung-Hoon Kim, Hyungjin Chung

机构 * EverEx University of Michigan(密歇根大学) ETH Zurich(苏黎世联邦理工学院) Yonsei University(延世大学)

专题命中 安全评测 :alignment(abstract)

Comments 18 pages, 7 figures, Project Page:https://hyelinnam.github.io/Cameo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24311 2025-10-07 cs.CV 50%

Towards Foundation Models for Cryo-ET Subtomogram Analysis

Runmin Jiang, Wanyue Feng, Yuntian Yang, Shriya Pingulkar, Hong Wang, Xi Xiao, Xiaoyu Cao, Genpei Zhang, Xiao Wang, Xiaolong Wu, Tianyang Wang, Yang Liu, Xingjian Li, Min Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Harvard University(哈佛大学) University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) Oak Ridge National Laboratory(橡树岭国家实验室) K. J. Somaiya College of Engineering(K.J. Somaiya 工程学院)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23054 2025-10-07 cs.CV 50%

Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis

Ruilang Wang, Shuotong Xu, Bowen Liu, Runlin Huang, Donglong Chen, Weifeng Su

机构 * Beijing Normal–Hong Kong Baptist University(北京师范大学-香港 Baptist大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏