arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-12 至 2025-11-12 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 11 篇

2511.08080 2025-11-12 cs.LG cs.AI 81%

Hierarchical Structure-Property Alignment for Data-Efficient Molecular Generation and Editing

Ziyu Fan, Zhijian Huang, Yahan Li, Xiaowen Hu, Siyuan Shen, Yunliang Wang, Zeyu Zhong, Shuhong Liu, Shuning Yang, Shangqian Wu, Min Wu, Lei Deng

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Central South University(中南大学) Institute for Infocomm Research Agency for Science, Technology and Research (A* STAR)(信息通信研究所(A* STAR))

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08402 2025-11-12 cs.CV cs.AI cs.LG 62%

Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation

Difei Gu, Yunhe Gao, Mu Zhou, Dimitris Metaxas

机构 * Rutgers University(新泽西罗格斯大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08052 2025-11-12 cs.AI cs.CL cs.SE 62%

Dual-Process Scaffold Reasoning for Enhancing LLM Code Debugging

Po-Chung Hsieh, Chin-Po Chen, Jeng-Lin Li, Ming-Ching Chang

机构 * National Taiwan University(国立台湾大学) AI Research Center, Inventec Corporation(Inventec公司人工智能研究中心) University at Albany, SUNY(纽约州立大学阿尔巴尼分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07699 2025-11-12 econ.GN cs.LG q-fin.EC 57%

Misaligned by Design: Incentive Failures in Machine Learning

David Autor, Andrew Caplin, Daniel Martin, Philip Marx

机构 * Massachusetts Institute of Technology, Google Technology and Society Fellows program, and NBER(麻省理工学院、谷歌技术与社会 fellows 程序及国家经济研究局) New York University and NBER(纽约大学及国家经济研究局) University of California, Santa Barbara(加州大学圣塔芭芭拉分校) Louisiana State University(路易斯安那州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17510 2025-11-12 cs.CL 57%

Large Language Models Do Multi-Label Classification Differently

Marcus Ma, Georgios Chochlakis, Niyantha Maruthu Pandiyan, Jesse Thomason, Shrikanth Narayanan

机构 * University of Southern California(南加州大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments To be published in the Main Conference Proceedings of EMNLP 2025, 24 pages, 16 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15056 2025-11-12 cs.LG cs.CV cs.GR 57%

ElastoGen: 4D Generative Elastodynamics

Yutao Feng, Yintong Shang, Xiang Feng, Lei Lan, Shandian Zhe, Tianjia Shao, Hongzhi Wu, Kun Zhou, Chenfanfu Jiang, Yin Yang

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17120 2025-11-12 cs.CL 57%

Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions

Dillon Plunkett, Adam Morris, Keerthi Reddy, Jorge Morales

机构 * Northeastern University(东北大学) Princeton University(普林斯顿大学) Independent Researcher(独立研究者)

专题命中 其他安全 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09205 2025-11-12 cs.MM cs.CL cs.IR cs.SD eess.AS 57%

Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model

Ali Vosoughi, Dimitra Emmanouilidou, Hannes Gamper

机构 * University of Rochester(罗切斯特大学) Microsoft Research(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted at EUSIPCO 2025 - 5 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11639 2025-11-12 cs.IR 50%

OneRec-Think: In-Text Reasoning for Generative Recommendation

Zhanyu Liu, Shiyao Wang, Xingmei Wang, Rongzhou Zhang, Jiaxin Deng, Honghui Bao, Jinghao Zhang, Wuchao Li, Pengfei Zheng, Xiangyu Wu, Yifei Hu, Qigen Hu, Xinchen Luo, Lejian Ren, Zixing Zhang, Qianqian Wang, Kuo Cai, Yunfan Wu, Hongtao Cheng, Zexuan Cheng, Lu Ren, Huanjie Wang, Yi Su, Ruiming Tang, Kun Gai, Guorui Zhou

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00598 2025-11-12 cs.CV 50%

DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation

Boyi Li, Ce Zhang, Richard M. Timmerman, Wenxuan Bao

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07598 2025-11-12 physics.app-ph 50%

Passive Acoustic Monitoring of Underwater Well Leakages with Machine Learning: A Review

Guanlin Zhu, Zechun Deng, Jiaxin Shen, Junchi Yang

专题命中 其他安全 :safety(abstract)

Comments 10 pages, 5 figures, 4 equations, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏