arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-17 至 2025-09-17 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2509.12750 2025-09-17 cs.CV 71%

What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment

Rishab Parthasarathy, Jasmine Collins, Cory Stephenson

专题命中 偏好对齐 :alignment(title)

Comments 7 pages, 9 figures, 3 tables; appendix 16 pages, 9 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10105 2025-09-17 cs.CV cs.CL 70%

VARCO-VISION-2.0 Technical Report

Young-rok Cha, Jeongho Ju, SunYoung Park, Jong-Hyeon Lee, Younghyun Yu, Youngjune Kim

机构 * NC AI

专题命中 偏好对齐 :alignment(abstract);safety(abstract);分类 cs.CL

Comments 19 pages, 1 figure, 14 tables. Technical report for VARCO-VISION-2.0, a Korean-English bilingual VLM in 14B and 1.7B variants. Key features: multi-image understanding, OCR with text localization, improved Korean capabilities

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12521 2025-09-17 cs.LG 57%

Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time

Yifan Lan, Yuanpu Cao, Weitong Zhang, Lu Lin, Jinghui Chen

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 偏好对齐 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12870 2025-09-17 eess.SP 50%

Towards personalized, precise and survey-free environment recognition: AI-enhanced sensor fusion without pre-deployment

Ruichen Wang, Zhikang Ni, Pengzhou Wang, Xiya Cao, Zhi Li, Bao Zhang

专题命中 偏好对齐 :RLHF(abstract)

Comments 5 pages, 7 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17711 2025-09-17 cs.SI 50%

Enhancing LLM-Based Social Bot via an Adversarial Learning Framework

Fanqi Kong, Xiaoyuan Zhang, Xinyu Chen, Yaodong Yang, Song-Chun Zhu, Xue Feng

专题命中 偏好对齐 :DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏