arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2511.02505 2025-11-20 cs.CV cs.AI 57%

ESA: Energy-Based Shot Assembly Optimization for Automatic Video Editing

Yaosen Chen, Wei Wang, Tianheng Zheng, Xuming Wen, Han Yang, Yanru Zhang

机构 * Sobey Media Intelligence Laboratory(索贝媒体智能实验室) University of Electronic Science and Technology of China(电子科学与技术大学) SiChuan University(四川大学) Qinghai Normal University(青海师范大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00275 2025-11-20 cs.CV cs.AI 57%

AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding

Md Asaduzzaman Jabin, Hanqi Jiang, Yiwei Li, Patrick Kaggwa, Eugene Douglass, Juliet N. Sekandi, Tianming Liu

机构 * University of Georgia(佐治亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: 7th International Workshop on Large Scale Holistic Video Understanding: Toward Video Foundation Models

Journal ref Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01293 2025-11-19 cs.CV cs.AI 57%

GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification

Ngoc Bui Lam Quang, Nam Le Nguyen Binh, Thanh-Huy Nguyen, Le Thien Phuc Nguyen, Quan Nguyen, Ulas Bagci

机构 * AI VIETNAM(AI越南) Carnegie Mellon University(卡内基梅隆大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) PTIT Northwestern University(西北大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Acccepted in MICCAI Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01916 2025-11-19 cs.RO cs.LG 57%

Generalizable and Fast Surrogates: Model Predictive Control of Articulated Soft Robots using Physics-Informed Neural Networks

Tim-Lukas Habich, Aran Mohammad, Simon F. G. Ehlers, Martin Bensch, Thomas Seel, Moritz Schappler

机构 * Leibniz University Hannover(莱布尼茨汉诺威大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments Accepted for publication in IEEE Transactions on Robotics (T-RO) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12511 2025-11-19 cs.CV cs.LG 57%

DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image Detection

Jialiang Shen, Jiyang Zheng, Yunqi Xue, Huajie Chen, Yu Yao, Hui Kang, Ruiqi Liu, Helin Gong, Yang Yang, Dadong Wang, Tongliang Liu

机构 * Sydney AI Center, The University of Sydney(悉尼人工智能中心,悉尼大学) CSIRO, Data61(澳大利亚联邦科学与工业研究组织、Data61) Shanghai Jiao Tong University(上海交通大学) City University of Macau(澳门城市大学) CASIA(中国科学院自动化研究所)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07991 2025-11-19 cs.AI 57%

VSPO: Validating Semantic Pitfalls in Ontology via LLM-Based CQ Generation

Hyojun Choi, Seokju Hwang, Kyong-Ho Lee

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted at AAAI 2026 oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12452 2025-11-18 cs.CV cs.CL 57%

DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions

Xiaoyu Lin, Aniket Ghorpade, Hansheng Zhu, Justin Qiu, Dea Rrozhani, Monica Lama, Mick Yang, Zixuan Bian, Ruohan Ren, Alan B. Hong, Jiatao Gu, Chris Callison-Burch

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18320 2025-11-18 cs.LG 57%

State of Health Estimation of Batteries Using a Time-Informed Dynamic Sequence-Inverted Transformer

Janak M. Patel, Milad Ramezankhani, Anirudh Deodhar, Dagnachew Birru

机构 * Applied Research, Quantiphi(Quantiphi应用研究)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19848 2025-11-18 cs.RO cs.AI 57%

Human-Centered AI and Autonomy in Robotics: Insights from a Bibliometric Study

Simona Casini, Pietro Ducange, Francesco Marcelloni, Lorenzo Pollini

机构 * Department of Information Engineering, University of Pisa(信息工程系,比萨大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments International Joint Conference on Neural Network 2025 - Accepted

Journal ref 10.1109/IJCNN64981.2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12167 2025-11-18 cond-mat.mtrl-sci cs.LG 57%

Rapid Machine Learning-Driven Detection of Pesticides and Dyes Using Raman Spectroscopy

Quach Thi Thai Binh, Thuan Phuoc, Xuan Hai, Thang Bach Phan, Vu Thi Hanh Thu, Nguyen Tuan Hung

机构 * Faculty of Physics and Physics Engineering, University of Science, Ho Chi Minh City 700000, Viet Nam(物理系和物理工程系,科学大学,胡志明市700000,越南) Center for Innovative Materials and Architectures (INOMAR)(创新材料与架构中心) Department of Materials Science and Engineering, National Taiwan University, Taipei 10617, Taiwan(材料科学与工程系,台湾国立大学,台北10617,台湾)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 25 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11885 2025-11-18 cs.DC cs.AI cs.DB 57%

Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs

Kausar Patherya, Ashutosh Dhekne, Francisco Romero

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 12 pages, 5 figures. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11780 2025-11-18 cs.CV cs.AI 57%

Image-POSER: Reflective RL for Multi-Expert Image Generation and Editing

Hossein Mohebbi, Mohammed Abdulrahman, Yanting Miao, Pascal Poupart, Suraj Kothawade

机构 * University of Waterloo(滑铁卢大学) Vector Institute(向量研究所) Google(谷歌)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10229 2025-11-14 cs.CL 57%

LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction Tuning

Yangfan Ye, Xiaocheng Feng, Xiachong Feng, Lei Huang, Weitao Ma, Qichen Hong, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments AAAI2026 Main Track Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10020 2025-11-14 cs.CV cs.AI 57%

Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation

Yuxin Jiang, Wei Luo, Hui Zhang, Qiyu Chen, Haiming Yao, Weiming Shen, Yunkang Cao

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09984 2025-11-14 cs.CL 57%

Language Drift in Multilingual Retrieval-Augmented Generation: Characterization and Decoding-Time Mitigation

Bo Li, Zhenghua Xu, Rui Xie

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments AAAI'26, Oral Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09973 2025-11-14 cs.CV cs.AI 57%

Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models

Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08238 2025-11-14 cs.CV cs.AI 57%

Remodeling Semantic Relationships in Vision-Language Fine-Tuning

Xiangyang Wu, Liu Liu, Baosheng Yu, Jiayan Qiu, Zhenwei Shi

机构 * Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) Nanyang Technological University(南洋理工大学) University of Leicester(莱斯特大学) School of Astronautics, Beihang University(北京航空航天大学航天学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05909 2025-11-14 cs.CL 57%

Beyond Perplexity: Let the Reader Select Retrieval Summaries via Spectrum Projection Score

Zhanghao Hu, Qinglin Zhu, Siya Qi, Yulan He, Hanqi Yan, Lin Gui

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted by AAAI 2026 Oral. Project link: https://zhanghao-aaai2026-sps.github.io/AAAI2026-SPS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22712 2025-11-14 cs.LG stat.ML 57%

Generalized Linear Mode Connectivity for Transformers

Alexander Theus, Alessandro Cabodi, Sotiris Anagnostidis, Antonio Orvieto, Sidak Pal Singh, Valentina Boeva

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted to NeurIPS 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09844 2025-11-14 cs.LG cs.PF 57%

Steering Pretrained Drafters during Speculative Decoding

Frédéric Berdoz, Peer Rheinboldt, Roger Wattenhofer

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09709 2025-11-14 cs.CL 57%

Contextual morphologically-guided tokenization for Latin encoder models

Marisa Hudspeth, Patrick J. Burns, Brendan O'Connor

机构 * Manning College of Information & Computer Sciences, University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校信息与计算机科学学院) Institute for the Study of the Ancient World, New York University(纽约大学古代世界研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23749 2025-11-13 astro-ph.IM cs.LG 57%

Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders

John F. Wu, Michael Walmsley

机构 * Space Telescope Science Institute(太空望远镜科学研究所) Johns Hopkins University(约翰霍普金斯大学) University of Toronto(多伦多大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Authors contributed equally to this work. Accepted to NeurIPS Machine Learning and the Physical Sciences Workshop. See trained model at https://huggingface.co/mwalmsley/euclid-rr2-mae, HuggingFace demo at https://huggingface.co/spaces/mwalmsley/euclid_masked_autoencoder, and code at https://github.com/jwuphysics/euclid-galaxy-morphology-saes

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08802 2025-11-13 cs.LG 57%

The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?

Denis Sutter, Julian Minder, Thomas Hofmann, Tiago Pimentel

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08593 2025-11-13 cs.CL 57%

Knowledge Graph Analysis of Legal Understanding and Violations in LLMs

Abha Jha, Abel Salinas, Fred Morstatter

专题命中 其他安全 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07699 2025-11-12 econ.GN cs.LG q-fin.EC 57%

Misaligned by Design: Incentive Failures in Machine Learning

David Autor, Andrew Caplin, Daniel Martin, Philip Marx

机构 * Massachusetts Institute of Technology, Google Technology and Society Fellows program, and NBER(麻省理工学院、谷歌技术与社会 fellows 程序及国家经济研究局) New York University and NBER(纽约大学及国家经济研究局) University of California, Santa Barbara(加州大学圣塔芭芭拉分校) Louisiana State University(路易斯安那州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17510 2025-11-12 cs.CL 57%

Large Language Models Do Multi-Label Classification Differently

Marcus Ma, Georgios Chochlakis, Niyantha Maruthu Pandiyan, Jesse Thomason, Shrikanth Narayanan

机构 * University of Southern California(南加州大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments To be published in the Main Conference Proceedings of EMNLP 2025, 24 pages, 16 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15056 2025-11-12 cs.LG cs.CV cs.GR 57%

ElastoGen: 4D Generative Elastodynamics

Yutao Feng, Yintong Shang, Xiang Feng, Lei Lan, Shandian Zhe, Tianjia Shao, Hongzhi Wu, Kun Zhou, Chenfanfu Jiang, Yin Yang

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17120 2025-11-12 cs.CL 57%

Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions

Dillon Plunkett, Adam Morris, Keerthi Reddy, Jorge Morales

机构 * Northeastern University(东北大学) Princeton University(普林斯顿大学) Independent Researcher(独立研究者)

专题命中 其他安全 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09205 2025-11-12 cs.MM cs.CL cs.IR cs.SD eess.AS 57%

Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model

Ali Vosoughi, Dimitra Emmanouilidou, Hannes Gamper

机构 * University of Rochester(罗切斯特大学) Microsoft Research(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted at EUSIPCO 2025 - 5 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07238 2025-11-11 cs.CV cs.AI 57%

Leveraging Text-Driven Semantic Variation for Robust OOD Segmentation

Seungheon Song, Jaekoo Lee

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 8 pages, 5 figure references, 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) submission

详情

展开后加载摘要…

URL PDF HTML 收藏