arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-26 至 2025-11-26 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7 篇

2511.20627 2025-11-26 cs.AI 79%

Fighting AI with AI: Leveraging Foundation Models for Assuring AI-Enabled Safety-Critical Systems

用AI对抗AI:利用基础模型确保AI赋能的安全关键系统

Anastasia Mavridou, Divya Gopinath, Corina S. Păsăreanu

机构 * KBR Inc.(KBR公司) NASA Ames(美国国家航空航天局阿姆斯研究中心)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

AI总结 本文提出利用AI技术解决安全关键系统中AI保证问题,通过REACT和SemaLens两个组件实现需求工程与感知系统验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20614 2025-11-26 cs.CV 78%

The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment

一致性批评者:通过参考引导的注意对齐纠正生成图像中的不一致

Ziheng Ouyang, Yiren Song, Yaoli Liu, Shihao Zhu, Qibin Hou, Ming-Ming Cheng, Mike Zheng Shou

机构 * VCIP, Nankai University Show Lab, National University of Singapore State Key Laboratory of CAD\&CG, Zhejiang University

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文提出ImageCritic,通过参考引导的注意力对齐方法纠正生成图像中的不一致问题,提升细粒度细节的一致性。

Comments Project page: https://ouyangziheng.github.io/ImageCritic-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19495 2025-11-26 cs.LG cs.AI 62%

A Systematic Study of Compression Ordering for Large Language Models

大语言模型压缩顺序的系统研究

Shivansh Chhawri, Rahul Mahadik, Suparna Rooj

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本研究系统探讨了大语言模型压缩技术的顺序对性能的影响,发现剪枝-知识蒸馏-量化顺序能实现3.68倍压缩并保持良好能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20344 2025-11-26 cs.CL 57%

The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models

类比的奇特案例:在大型语言模型中探讨类比推理

Taewhoo Lee, Minju Song, Chanwoong Yoon, Jungwoo Park, Jaewoo Kang

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本研究探讨了大型语言模型在类比推理中的能力,发现其在编码和应用高层关系概念方面表现出有限但新兴的能力,揭示了与人类认知的相似性和差距。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19471 2025-11-26 eess.IV cs.AI cs.CV 57%

Not Quite Anything: Overcoming SAMs Limitations for 3D Medical Imaging

并非一切:克服SAMs在3D医学影像中的局限性

Keith Moore

机构 * Deptartment of Biomedical Data Science Stanford University(生物医学数据科学系 斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出了一种无需微调基础模型的组合性替代方案,通过将基础模型输出作为额外输入通道来提高3D医学影像分割的准确性与鲁棒性。

Comments Preprint; Paper accepted at AIAS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20186 2025-11-26 cs.CV 50%

Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis

Exo2EgoSyn: 解锁用于外视图到内视图视频合成的基础视频生成模型

Mohammad Mahdi, Yuqian Fu, Nedko Savov, Jiancheng Pan, Danda Pani Paudel, Luc Van Gool

专题命中 其他安全 :alignment(abstract)

AI总结 Exo2EgoSyn通过三个模块实现从第三人称视角生成高保真内视图视频,提升跨视角视频合成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17964 2025-11-26 cs.CV 50%

X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification

X-ReID:基于视频的可见-红外人重识别中的多粒度信息交互

Chenyang Yu, Xuehu Liu, Pingping Zhang, Huchuan Lu

专题命中 其他安全 :alignment(abstract)

AI总结 X-ReID通过多粒度信息交互和跨模态特征学习,提升视频中可见-红外人重识别的性能。

Comments Accepted by AAAI2026. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏