Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs
机构 * Dataplicada
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Dataplicada
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY
Comments To appear in NeurIPS 2025
机构 * Koguan School of Law, China Institute for Smart Justice, School of Computer Science, Shanghai Jiao Tong University(柯 guar 法学院、中国智能正义研究院、计算机科学学院、上海交通大学)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
专题命中 AI治理与伦理 :alignment(abstract)
Comments Accepted by AAAI2026
机构 * University of Trento(特伦托大学)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Published in Transactions on Machine Learning Research (TMLR)
机构 * Global Technical Service (GTS) Huawei Technologies Co., Ltd(华为技术有限公司全球技术服务部)
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG
Comments Accepted for Oral Presentation at the 40th AAAI Conference on Artificial Intelligence (AAAI-26), Main Technical Track
机构 * University of the Bundeswehr Munich(联邦国防军慕尼黑大学)
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL
Comments Accepted by AAAI 2026 as an Oral Presentation (13 pages, 7 figures, 7 tables)
Journal ref AAAI2026
机构 * Stanford University(斯坦福大学) ; Mathematical Medicine Group(数学医学组) ; Department of Neurosurgery(神经外科系) ; Physician-Scientist Training Program(医师科学家培训计划)
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.CY
机构 * University of New South Wales(新南威尔士大学)
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG
机构 * The University of Tokyo(东京大学) ; Amazon.com, Inc.(亚马逊公司)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Ontario Tech University(安大略技术大学) ; Stanford University(斯坦福大学) ; SAS Posos ; University of Toronto(多伦多大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
Comments Accepted to IJCNLP-AACL 2025
机构 * College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
Comments This paper was accepted by IEEE BIBM 2025 conference
专题命中 其他安全 :alignment(abstract);分类 cs.CY、cs.LG
机构 * South China University of Technology(华南理工大学) ; The University of Hong Kong(香港大学) ; City University of Hong Kong(城市大学)
专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.CY
Comments https://github.com/HKUDS/Awesome-LLM4Urban-Papers
Journal ref ACM TIST, 2025
机构 * College of Computer Science and Artificial Intelligence(计算机科学与人工智能学院) ; Fudan University(复旦大学) ; Institute of Modern Languages and Linguistics(现代语言与语言学研究所) ; Shanghai Key Laboratory of Intelligent Information Processing(上海智能信息处理重点实验室) ; Research Institute of Intelligent Complex Systems(智能复杂系统研究所) ; Institute of Trustworthy Embodied Artificial Intelligence(可信具身人工智能研究所) ; State Key Laboratory of Genetics and Development of Complex Phenotypes(复杂表型遗传与发育国家重点实验室)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
Comments 66 pages. Accepted manuscript. Final version published in Proceedings of the National Academy of Sciences (PNAS): https://www.pnas.org/doi/10.1073/pnas.2512514122
Journal ref Proceedings of the National Academy of Sciences, U.S.A., 122 (44) e2512514122 (2025)
专题命中 其他安全 :safety(abstract);分类 cs.AI
Comments 8 pages, 5 figure references, 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) submission
机构 * University of Applied Sciences Ruhr West(鲁尔西部应用科学大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
机构 * York University(约克大学) ; Microsoft Research(微软研究院)
专题命中 其他安全 :safety(abstract);分类 cs.AI
Comments This work has been submitted to the IEEE for possible publication
专题命中 其他安全 :safety(abstract);分类 cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.AI
机构 * Algoverse AI Research(Algoverse AI研究)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
Comments Multi-Turn Interactions in Large Language Models (MTI-LLM) Workshop at NeurIPS 2025
专题命中 其他安全 :alignment(abstract);分类 cs.CL
机构 * School of Computer Science, University of Bristol(布里斯托大学计算机科学学院)
专题命中 其他安全 :alignment(abstract)
Comments 14 pages, 6 figures
专题命中 其他安全 :alignment(abstract)
机构 * Mercedes-Benz AG(梅赛德斯-奔驰集团) ; University of Tübingen(图宾根大学) ; Tübingen AI Center(图宾根人工智能中心) ; University of Bonn(波恩大学) ; RPL ; KTH Royal Institute of Technology(皇家理工学院) ; TU Berlin(柏林技术大学)
专题命中 其他安全 :alignment(abstract)
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; The Chinese University of Hong Kong(香港中文大学) ; Tsinghua University(清华大学) ; SenseTime Research(商汤科技研究院) ; Shanghai Jiao Tong University(上海交通大学) ; Peking University(北京大学) ; PengCheng Laboratory(鹏城实验室) ; Chongqing University(重庆大学)
专题命中 其他安全 :alignment(abstract)
Comments The code and weights are open-sourced. Project page: https://curryx-001.github.io/LangBridge.github.io/
专题命中 其他安全 :alignment(abstract)