arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-26 至 2025-11-26 共收录 53 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 21 篇

2511.19480 2025-11-26 cs.LG cs.AI 62%

Exploiting the Experts: Unauthorized Compression in MoE-LLMs

利用专家:MoE-LLMs中的未经授权压缩

Pinaki Prasad Guha Neogi, Ahmad Mohammadshirazi, Dheeraj Kulshrestha, Rajiv Ramnath

机构 * Ohio State University(俄亥俄州立大学) Flairsoft

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了MoE-LLMs在特定任务中的可剪枝性,提出专家归因框架和防御策略以防止未经授权的压缩和微调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20650 2025-11-26 cs.CV cs.AI 57%

MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities

MedROV:面向跨多种医学影像模态的实时开放词汇检测

Tooba Tehreem Sheikh, Jean Lahoud, Rao Muhammad Anwer, Fahad Shahbaz Khan, Salman Khan, Hisham Cholakkal

机构 * Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 MedROV是首个实时开放词汇医学影像检测模型,通过大规模数据集和伪标签策略提升检测性能,实现40 mAP50的提升并达到70 FPS的实时处理速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20526 2025-11-26 cs.AI 57%

Assessing LLMs' Performance: Insights from the Chinese Pharmacist Exam

评估大语言模型的性能:来自中国药师资格考试的洞察

Xinran Wang, Boran Zhu, Shujuan Zhou, Ziwen Long, Dehua Zhou, Shu Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本研究比较了DeepSeek-R1和ChatGPT-4o在药师资格考试中的表现,发现DeepSeek-R1在准确性上显著优于后者,强调了领域特定模型在评估中的重要性及人类监督的必要性。

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19952 2025-11-26 cs.LG 57%

Hierarchical Spatio-Temporal Attention Network with Adaptive Risk-Aware Decision for Forward Collision Warning in Complex Scenarios

具有自适应风险感知决策的分层时空注意力网络用于复杂场景的前方碰撞预警

Haoran Hu, Junren Shi, Shuo Jiang, Kun Cheng, Xia Yang, Changhao Piao

机构 * organization= School of Automation \& School of Industrial Internet, Chongqing University of Posts organization= Platform Technology Development Department, AVATR Technology Co. LTD. , city= Chongqing , postcode= 400000 , country= China organization= School of Vehicle Mobility, Tsinghua University , city= Beijing , postcode= 100084 , country= China

专题命中 安全评测 :safety(abstract);分类 cs.LG

AI总结 本文提出一种结合分层时空注意力网络和动态风险阈值调整算法的前方碰撞预警框架,通过高效模型和自适应机制提升复杂场景下的预警精度与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19657 2025-11-26 cs.LG 57%

Structured Noise Modeling for Enhanced Time-Series Forecasting

结构噪声建模以提升时间序列预测

Sepideh Koohfar

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

AI总结 本文提出结构化噪声建模框架,通过模糊-去噪机制提升时间序列预测的准确性与稳定性,适用于能源、基础设施等时间敏感领域。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19654 2025-11-26 cs.CR cs.AI 57%

Accuracy and Efficiency Trade-Offs in LLM-Based Malware Detection and Explanation: A Comparative Study of Parameter Tuning vs. Full Fine-Tuning

在基于大语言模型的恶意软件检测与解释中的准确性与效率权衡:参数调优与全微调的比较研究

Stephen C. Gravereaux, Sheikh Rabiul Islam

机构 * Department of Cybersecurity University at Albany - State University of New York Albany, NY, USA(网络安全系 美国纽约州立大学阿尔巴尼分校)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 本研究比较了参数调优与全微调在恶意软件检测中的性能,发现LoRA在保持解释质量的同时显著提升了效率。

Comments Accepted in IEEE Big Data 2025

Journal ref IEEE Big Data 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01265 2025-11-26 cs.CL 57%

AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs

AraFinNews: 基于领域适应大语言模型的阿拉伯语金融摘要

Mo El-Haj, Paul Rayson

机构 * School of Computing(计算学院) Lancaster University(兰卡斯特大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 AraFinNews通过领域适应大语言模型提升阿拉伯语金融文本摘要的连贯性和准确性。

Comments 9 pages

Journal ref IEEE BigData 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15097 2025-11-26 cs.CR cs.AI 57%

MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm

MAIF:基于 artifacts 的代理范式强化 AI 可信度与溯源

Vineeth Sai Narajala, Manish Bhatt, Idan Habler, Ronald F. Del Rosario, Ads Dawson

机构 * Security Researcher Cisco(安全研究员 希思科) Researcher OWASP/Project Kuiper Security(研究员 OWASP/项目 Kuiper 安全) Adversarial AI Security Research Cisco(对抗性AI安全研究 希思科)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 MAIF 通过以 artifacts 为中心的代理范式,解决 AI 的可信度、安全性和问责问题,实现数据的主动信任执行和高效处理。

Comments 7 Pages, 2 Figures, 6 Tables, Repo: https://github.com/vineethsai/maifscratch-1, Added additional Author and fixed Citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20550 2025-11-26 cs.LO 50%

Verifying Numerical Methods with Isabelle/HOL

用Isabelle/HOL验证数值方法

Dustin Bryant, Jonathan Julian Huerta y Munive, Simon Foster

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出基于ITrees的Isabelle/HOL框架,用于验证数值方法,通过形式化规范和自动证明方法,实现从形式化到可执行代码的端到端验证流程。

Comments 30 pages, 30 listings, for accompanying formalisation, see https://zenodo.org/records/17679526

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20218 2025-11-26 cs.CV 50%

Text-guided Controllable Diffusion for Realistic Camouflage Images Generation

基于文本引导的可控扩散生成逼真伪装图像

Yuhang Qian, Haiyan Chen, Wentong Li, Ningzhong Liu, Jie Qin

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出CT-CIG方法,通过文本引导和可控扩散生成逼真且逻辑合理的伪装图像,利用VLM和FIRM模块提升伪装图像质量。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09436 2025-11-26 cs.CV 50%

Scaling up self-supervised learning for improved surgical foundation models

提升自监督学习以改进手术基础模型

Tim J. M. Jaspers, Ronald L. P. D. de Jong, Yiping Li, Carolus H. J. Kusters, Franciscus H. A. Bakker, Romy C. van Jaarsveld, Gino M. Kuiper, Richard van Hillegersberg, Jelle P. Ruurda, Willem M. Brinkman, Josien P. W. Pluim, Peter H. N. de With, Marcel Breeuwer, Yasmina Al Khalil, Fons van der Sommen

机构 * Department of Electrical Engineering, Video Coding \& Architectures, Eindhoven University of Technology, Eindhoven, The Netherlands Department of Biomedical Engineering, Medical Image Analysis, Eindhoven University of Technology, Eindhoven, The Netherlands Department of Surgery, University Medical Center Utrecht, Utrecht, The Netherlands Department of Oncological Urology, University Medical Center Utrecht, Utrecht, The Netherlands Department of Urology, Catharina Hospital, Eindhoven, The Netherlands

专题命中 安全评测 :safety(abstract)

AI总结 本研究提出SurgeNetXL,通过大规模预训练提升手术计算机视觉性能,实现多个任务上的显著改进。

Journal ref Medical Image Analysis, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19526 2025-11-26 cs.CV 50%

Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models

感知分类:评估和引导视觉语言模型中的层次场景推理

Jonathan Lee, Xingrui Wang, Jiawei Peng, Luoxin Ye, Zehan Zheng, Tiezheng Zhang, Tao Wang, Wufei Ma, Siyi Chen, Yu-Cheng Chou, Prakhar Kaushik, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出感知分类基准测试,旨在评估和引导视觉语言模型在层次场景推理中的能力,揭示模型在属性驱动推理上的不足,并展示通过上下文示例提升性能的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 4 篇

2511.19749 2025-11-26 cs.AI 79%

Scaling Item-to-Standard Alignment with Large Language Models: Accuracy, Limits, and Solutions

基于大语言模型的项目-标准对齐扩展:准确性、限制与解决方案

Farzan Karimi-Malekabadi, Pooya Razavi, Sonya Powers

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

AI总结 本研究探讨了大语言模型在提升项目-标准对齐效率方面的应用,通过实验发现LLMs在识别不一致项目和筛选候选技能方面表现优异,结合过滤策略可显著减少人工工作量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08087 2025-11-26 cs.CR 75%

Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks

保障大语言模型:应对偏见、虚假信息和提示攻击

Benji Peng, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Junyu Liu, Xinyuan Song, Qian Niu

专题命中 AI治理与伦理 :jailbreak(abstract);red teaming(abstract);prompt injection(abstract)

AI总结 本文探讨了大语言模型在偏见、虚假信息和提示攻击方面的安全问题,分析了偏见缓解策略、内容检测机制及防御措施,强调了对LLM安全领域进一步研究的重要性。

Comments 17 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19979 2025-11-26 cs.IR 67%

The 2nd Workshop on Human-Centered Recommender Systems

人类中心推荐系统研讨会第二届

Kaike Zhang, Jiakai Tang, Du Su, Shuchang Liu, Julian McAuley, Lina Yao, Qi Cao, Yue Feng, Fei Sun

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract)

AI总结 该研讨会旨在推动推荐系统从优化参与度向设计真正理解、参与和惠及人类的系统转变,探讨如何整合人类价值观以提升推荐系统的社会责任感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19886 2025-11-26 cs.CR cs.CV 50%

Frequency Bias Matters: Diving into Robust and Generalized Deep Image Forgery Detection

频率偏差至关重要:深入探讨鲁棒且通用的深度图像伪造检测

Chi Liu, Tianqing Zhu, Wanlei Zhou, Wei Zhao

机构 * Faculty of Data Science, City University of Macau(数据科学学院,澳门城市大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院)

专题命中 AI治理与伦理 :alignment(abstract)

AI总结 本文从频率视角分析深度图像伪造检测器的通用性和鲁棒性问题,提出两步频率对齐方法以提升检测可靠性并增强反伪造能力。

Comments Accepted for publication in IEEE Transactions on Dependable and Secure Computing

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 7 篇

2511.20627 2025-11-26 cs.AI 79%

Fighting AI with AI: Leveraging Foundation Models for Assuring AI-Enabled Safety-Critical Systems

用AI对抗AI:利用基础模型确保AI赋能的安全关键系统

Anastasia Mavridou, Divya Gopinath, Corina S. Păsăreanu

机构 * KBR Inc.(KBR公司) NASA Ames(美国国家航空航天局阿姆斯研究中心)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

AI总结 本文提出利用AI技术解决安全关键系统中AI保证问题,通过REACT和SemaLens两个组件实现需求工程与感知系统验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20614 2025-11-26 cs.CV 78%

The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment

一致性批评者:通过参考引导的注意对齐纠正生成图像中的不一致

Ziheng Ouyang, Yiren Song, Yaoli Liu, Shihao Zhu, Qibin Hou, Ming-Ming Cheng, Mike Zheng Shou

机构 * VCIP, Nankai University Show Lab, National University of Singapore State Key Laboratory of CAD\&CG, Zhejiang University

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文提出ImageCritic,通过参考引导的注意力对齐方法纠正生成图像中的不一致问题,提升细粒度细节的一致性。

Comments Project page: https://ouyangziheng.github.io/ImageCritic-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19495 2025-11-26 cs.LG cs.AI 62%

A Systematic Study of Compression Ordering for Large Language Models

大语言模型压缩顺序的系统研究

Shivansh Chhawri, Rahul Mahadik, Suparna Rooj

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本研究系统探讨了大语言模型压缩技术的顺序对性能的影响,发现剪枝-知识蒸馏-量化顺序能实现3.68倍压缩并保持良好能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20344 2025-11-26 cs.CL 57%

The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models

类比的奇特案例:在大型语言模型中探讨类比推理

Taewhoo Lee, Minju Song, Chanwoong Yoon, Jungwoo Park, Jaewoo Kang

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本研究探讨了大型语言模型在类比推理中的能力,发现其在编码和应用高层关系概念方面表现出有限但新兴的能力,揭示了与人类认知的相似性和差距。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19471 2025-11-26 eess.IV cs.AI cs.CV 57%

Not Quite Anything: Overcoming SAMs Limitations for 3D Medical Imaging

并非一切:克服SAMs在3D医学影像中的局限性

Keith Moore

机构 * Deptartment of Biomedical Data Science Stanford University(生物医学数据科学系 斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出了一种无需微调基础模型的组合性替代方案,通过将基础模型输出作为额外输入通道来提高3D医学影像分割的准确性与鲁棒性。

Comments Preprint; Paper accepted at AIAS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20186 2025-11-26 cs.CV 50%

Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis

Exo2EgoSyn: 解锁用于外视图到内视图视频合成的基础视频生成模型

Mohammad Mahdi, Yuqian Fu, Nedko Savov, Jiancheng Pan, Danda Pani Paudel, Luc Van Gool

专题命中 其他安全 :alignment(abstract)

AI总结 Exo2EgoSyn通过三个模块实现从第三人称视角生成高保真内视图视频,提升跨视角视频合成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17964 2025-11-26 cs.CV 50%

X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification

X-ReID:基于视频的可见-红外人重识别中的多粒度信息交互

Chenyang Yu, Xuehu Liu, Pingping Zhang, Huchuan Lu

专题命中 其他安全 :alignment(abstract)

AI总结 X-ReID通过多粒度信息交互和跨模态特征学习,提升视频中可见-红外人重识别的性能。

Comments Accepted by AAAI2026. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏