arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-22 至 2025-09-22 共收录 16 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 16 篇

2503.12613 2025-09-22 cs.HC cs.AI cs.CY cs.MA 81%

Negotiative Alignment: Embracing Disagreement to Achieve Fairer Outcomes -- Insights from Urban Studies

Rashid Mushkani, Hugo Berard, Shin Koseki

机构 * Université de Montréal(蒙特利尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments 16 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15940 2025-09-22 cs.DC 71%

Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs

Guoliang He, Youhe Jiang, Wencong Xiao, Kaihua Jiang, Shuguang Wang, Jun Wang, Zixian Du, Zhuo Jiang, Xinlei Zhang, Binhang Yuan, Eiko Yoneki

专题命中 其他安全 :alignment(title)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15243 2025-09-22 cs.CV 67%

Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models

Muhammad Imran, Yugyung Lee

机构 * Computer Science, School of Science and Engineering, University of Missouri - Kansas City(计算机科学系,科学与工程学院,密苏里大学-堪萨斯城分校)

专题命中 其他安全 :alignment(abstract);safety(abstract)

Comments 8 pages, 6 figures, 3 tables

Journal ref Non-Archival track - The First Workshop on Multimodal Knowledge and Language Modeling IJCAI 2025 Workshop, August 16, 2025 IJCAI 2025 Workshop, August 16, 2025 Room 516B, Palais des congrès, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15789 2025-09-22 cs.CL cs.LG 62%

UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations

Qiuyang Lu, Fangjian Shen, Zhengkai Tang, Qiang Liu, Hexuan Cheng, Hui Liu, Wushao Wen

机构 * United Nations(联合国)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments 5 pages, 1 figure, submitted to ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15975 2025-09-22 cs.CL cs.AI 62%

Sparsity May Be All You Need: Sparse Random Parameter Adaptation

Jesus Rios, Pierre Dognin, Ronny Luss, Karthikeyan N. Ramamurthy

机构 * IBM Research(IBM研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18848 2025-09-22 cs.LG cs.AI 62%

Two Is Better Than One: Aligned Representation Pairs for Anomaly Detection

Alain Ryser, Thomas M. Sutter, Alexander Marx, Julia E. Vogt

机构 * Department of Computer Science ETH Zurich(计算机科学系,苏黎世联邦理工学院) Research Center Trustworthy Data Science and Security of the University Alliance Ruhr(可信数据科学与安全的鲁尔大学联盟研究中心) Department of Statistics TU Dortmund University(统计学系,多特蒙德技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Published in Transactions on Machine Learning Research (TMLR) https://openreview.net/forum?id=Bt0zdsnWYc

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15474 2025-09-22 cs.CL cs.AI 62%

Subjective Behaviors and Preferences in LLM: Language of Browsing

Sai Sundaresan, Harshita Chopra, Atanu R. Sinha, Koustava Goswami, Nagasai Saketh Naidu, Raghav Karan, N Anushka

机构 * Adobe Research(Adobe研究院) University of Washington, Seattle(华盛顿大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15740 2025-09-22 cs.LG 57%

Incremental Multistep Forecasting of Battery Degradation Using Pseudo Targets

Jonathan Adam Rico, Nagarajan Raghavan, Senthilnath Jayavelu

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments The published version of this preprint can be accessed at https://ieeexplore.ieee.org/abstract/document/10874675

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15667 2025-09-22 cs.CL cs.SD eess.AS 57%

VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion

Dimitrios Damianos, Leon Voukoutis, Georgios Paraskevopoulos, Vassilis Katsouros

机构 * Institute for Speech and Language Processing, Athena Research Center, Greece(语音与语言处理研究所,亚特兰蒂斯研究中心,希腊)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05244 2025-09-22 cs.CV cs.AI 57%

RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding

Tianchen Fang, Guiru Liu

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Upon further review, we identified that our dataset requires optimization to ensure research reliability and accuracy. Additionally, considering the target journal's latest submission policies, we believe comprehensive manuscript revisions are necessary

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12260 2025-09-22 cs.CL 57%

Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese

Yikang Liu, Wanyang Zhang, Yiming Wang, Jialong Tang, Pei Zhang, Baosong Yang, Fei Huang, Rui Wang, Hai Hu

机构 * Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) Tongyi Lab(通义实验室) City University of Hong Kong(香港城市大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14031 2025-09-22 cs.NE cs.LG 57%

Modeling the Human Visual System: Comparative Insights from Response-Optimized and Task-Optimized Vision Models, Language Models, and different Readout Mechanisms

Shreya Saha, Ishaan Chadha, Meenakshi Khosla

机构 * Electrical and Computer Engineering University of California, San Diego(电气与计算机工程大学加州大学圣地亚哥分校) Halıcıoğlu Data Science Institute University of California, San Diego(Halıcıoğlu数据科学研究所大学加州大学圣地亚哥分校) Department of Cognitive Science, Department of Computer Science and Engineering University of California, San Diego(认知科学系计算机科学与工程系大学加州大学圣地亚哥分校)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12812 2025-09-22 cs.RO cs.AI 57%

Towards Interactive and Learnable Cooperative Driving Automation: a Large Language Model-Driven Decision-Making Framework

Shiyu Fang, Jiaqi Liu, Mingyu Ding, Yiming Cui, Chen Lv, Peng Hang, Jian Sun

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07733 2025-09-22 cs.CV cs.AI 57%

Beyond Pixels: Enhancing LIME with Hierarchical Features and Segmentation Foundation Models

Patrick Knab, Sascha Marton, Christian Bartelt

机构 * Clausthal University of Technology(Clausthal 技术大学) University of Mannheim(曼海姆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments ECAI 2025 - Main Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16006 2025-09-22 cs.RO cs.HC 50%

Defining and Monitoring Complex Robot Activities via LLMs and Symbolic Reasoning

Francesco Argenziano, Elena Umili, Francesco Leotta, Daniele Nardi

机构 * Department of Computer, Automation and Management Engineering(计算机、自动化与管理工程系)

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15737 2025-09-22 cs.RO cs.SY eess.SY 50%

SMART: Scalable Multi-Agent Reasoning and Trajectory Planning in Dense Environments

Heye Huang, Yibin Yang, Wang Chen, Tiantian Chen, Xiaopeng Li, Sikai Chen

机构 * Department of Civil and Environmental Engineering University of Wisconsin-Madison(威斯康星大学麦迪逊分校土木与环境工程系) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动性学院) Department of Civil Engineering, The University of Hong Kong(香港大学土木工程系)

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏