arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-22 至 2026-01-22 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 11 篇

2601.14728 2026-01-22 eess.AS cs.AI cs.CL cs.LG cs.SD 82%

AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering

AQAScore: 通过音频问答评估文本到音频生成中的语义对齐

Chun-Yi Kuan, Kai-Wei Chang, Hung-yi Lee

机构 * National Taiwan University(国立台湾大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 AQAScore通过音频问答评估文本到音频生成的语义对齐,利用大型语言模型的推理能力,有效捕捉语义不一致性。

Comments Manuscript in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15267 2026-01-22 cs.CY cs.AI cs.CL 67%

Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions

评估大型语言模型在法律应用中的表现:挑战、方法与未来方向

Yiran Hu, Huanghai Liu, Chong Wang, Kunran Li, Tien-Hsuan Wu, Haitao Li, Xinran Xu, Siqing Huo, Weihang Su, Ning Zheng, Siyuan Zheng, Qingyao Ai, Yun Liu, Renjun Bian, Yiqun Liu, Charles L. A. Clarke, Weixing Shen, Ben Kao

机构 * Tsinghua University(清华大学) The University of Hong Kong(香港大学) University of Waterloo(多伦多大学) Shanghai Jiaotong University(上海交通大学) Peking University(北京大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文探讨了大型语言模型在法律应用中的评估挑战与方法,分析了现有评估框架的局限性,并提出了未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14620 2026-01-22 eess.AS cs.AI cs.LG cs.SD 62%

Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models

缩放歧义:通过音频-语言模型增强语音情感识别中的人工标注

Wenda Zhang, Hongyu Jin, Siyi Wang, Zhiqiang Wei, Ting Dang

机构 * University of Melbourne(墨尔本大学) Xi'an Jiaotong University(西安交通大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文通过音频-语言模型生成合成标注,提升语音情感识别中模糊情感的真实分布可靠性,实验表明合成标注在低歧义区域效果显著,但高歧义情况需进一步优化。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14269 2026-01-22 cs.CL cs.AI 62%

The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues

支持的缓慢漂移:多轮心理健康大语言模型对话中的边界失败

Youyou Cheng, Zhuangwei Kang, Kerry Jiang, Chenyu Sun, Qiyang Pan

机构 * University of Incarnate Word School of Osteopathic Medicine(incarnate Word 学校医学部) Independent Researcher(独立研究者) Mayo Clinic(梅奥诊所)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出多轮压力测试框架,揭示LLMs在长对话中因安慰和同理心尝试导致的安全边界逐步侵犯问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14235 2026-01-22 astro-ph.IM astro-ph.CO cs.AI cs.LG stat.ML 62%

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

人工智能/机器学习在Rubin LSST暗能量科学合作中的机遇

LSST Dark Energy Science Collaboration, Eric Aubourg, Camille Avestruz, Matthew R. Becker, Biswajit Biswas, Rahul Biswas, Boris Bolliet, Adam S. Bolton, Clecio R. Bom, Raphaël Bonnet-Guerrini, Alexandre Boucaud, Jean-Eric Campagne, Chihway Chang, Aleksandra Ćiprijanović, Johann Cohen-Tanugi, Michael W. Coughlin, John Franklin Crenshaw, Juan C. Cuevas-Tello, Juan de Vicente, Seth W. Digel, Steven Dillmann, Mariano Javier de León Dominguez Romero, Alex Drlica-Wagner, Sydney Erickson, Alexander T. Gagliano, Christos Georgiou, Aritra Ghosh, Matthew Grayling, Kirill A. Grishin, Alan Heavens, Lindsay R. House, Mustapha Ishak, Wassim Kabalan, Arun Kannawadi, François Lanusse, C. Danielle Leonard, Pierre-François Léget, Michelle Lochner, Yao-Yuan Mao, Peter Melchior, Grant Merz, Martin Millon, Anais Möller, Gautham Narayan, Yuuki Omori, Hiranya Peiris, Laurence Perreault-Levasseur, Andrés A. Plazas Malagón, Nesar Ramachandra, Benjamin Remy, Cécile Roucelle, Jaime Ruiz-Zapatero, Stefan Schuldt, Ignacio Sevilla-Noarbe, Ved G. Shah, Tjitske Starkenburg, Stephen Thorp, Laura Toribio San Cipriano, Tilman Tröster, Roberto Trotta, Padma Venkatraman, Amanda Wasserman, Tim White, Justine Zeghal, Tianqing Zhang, Yuanyuan Zhang

机构 * Université Paris Cité, CNRS, CEA, Astroparticule et Cosmologie, F-75013 Paris, France Department of Physics, University of Michigan, Ann Arbor, MI 48109, USA Leinweber Institute of Theoretical Physics, University of Michigan, Ann Arbor, MI 48109, USA Argonne National Laboratory, 9700 South Cass Avenue, Lemont, IL 60439, USA Cavendish Astrophysics, University of Cambridge, Madingley Road, Cambridge CB3 0HA, UK Kavli Institute for Cosmology, University of Cambridge, Madingley Road, Cambridge CB3 0HA, UK SLAC National Accelerator Laboratory, Menlo Park, CA 94025, USA Department of Computer Science, University of Milan, Milan, Italy Université Paris Cité, CNRS, Astroparticule et Cosmologie, F-75013 Paris, France Université Paris-Saclay, CNRS/IN2P3, IJCLab, 91405 Orsay, France Department of Astronomy Astrophysics, University of Chicago, Chicago, IL 60637, USA Kavli Institute for Cosmological Physics, University of Chicago, Chicago, IL 60637, USA NSF-Simons AI Institute for the Sky (SkAI), 172 E. Chestnut St., Chicago, IL 60611, USA Fermi National Accelerator Laboratory, P.O. Box 500, Batavia, IL 60510, USA Universit\'e Clermont-Auvergne, CNRS, LPCA, 63000 Clermont-Ferrand, France Kavli Institute for Particle Astrophysics Cosmology, Stanford University, Stanford, CA 94305, USA Department of Physics, Stanford University, 382 Via Pueblo Mall, Stanford, CA 94305, USA Engineering Faculty, Universidad Autonoma de San Luis Potosi, Zona Universitaria, San Luis Potosi, 78290, Mexico Stanford Artificial Intelligence Laboratory, Stanford University, Stanford, CA 94305, USA Kavli Institute of Cosmological Physics, University of Chicago, Chicago, IL 60637, USA The NSF AI Institute for Artificial Intelligence Center for Astrophysics Harvard \& Smithsonian, 60 Garden Street, Cambridge, MA 02138, USA Department of Physics Kavli Institute for Astrophysics Space Research, Massachusetts Institute of Technology, Cambridge, MA 02139, USA Institut de Física d'Altes Energies (IFAE), The Barcelona Institute of Science Institute of Astronomy Kavli Institute for Cosmology, University of Cambridge, Madingley Road, Cambridge, CB3 0HA, UK Imperial Centre for Inference Cosmology (ICIC), Imperial College London, Blackett Laboratory, Prince Consort Road, London SW7 2AZ, UK Data Science Institute, The University of Chicago, Chicago, IL 60615, USA Department of Physics, The University of Texas at Dallas, Richardson, TX 75080, USA Department of Physics, Duke University, Durham, NC 27708, USA Université Paris-Saclay, Université Paris Cité, CEA, CNRS, AIM, F-91191 Gif-sur-Yvette, France School of Mathematics, Statistics Physics, Newcastle University, Newcastle upon Tyne, NE1 7RU, United Kingdom Department of Astrophysical Sciences, Princeton University, Princeton, NJ 08544, USA Astronomy, University of the Western Cape, Bellville, Cape Town, 7535, South Africa Astronomy, University of Utah, Salt Lake City, UT 84112, USA Department of Astrophysical Sciences, Princeton University, Peyton Hall, Princeton, NJ 08544, USA Department of Astronomy, University of Illinois Urbana Champaign, 1002 W. Green St., Urbana, IL, 61801, USA Institute for Particle Physics Astrophysics, ETH Zürich, Wolfgang-Pauli-Strasse 27, CH-8093 Zurich, Switzerland Swinburne University of Technology, Hawthorn, Victoria 3122, Australia Ciela - Montr\'eal Institute for Astrophysical Data Analysis Mila - Quebec Artificial Intelligence Institute, Montréal, QC H2S 3H1, Canada Advanced Research Computing Centre, University College London, 90 High Holborn, London WC1V 6LJ, UK Finnish Centre for Astronomy with ESO (FINCA), University of Turku, FI-20014 Turku, Finland Department of Physics, P.O. Box 64, University of Helsinki, FI-00014 Helsinki, Finland Astronomy, Northwestern University, Evanston, IL, USA Center for Interdisciplinary Exploration Research in Astrophysics, Northwestern University, Evanston, IL, USA Scientific Data Science, International School for Advanced Study, Via Bonomea 265, I-34136 Trieste, Italy Department of Statistics, University of Michigan, Ann Arbor, MI 48109, USA PITT PACC, University of Pittsburgh, Pittsburgh, PA 15260, USA NSF NOIRLab, 950 N. Cherry Ave., Tucson, AZ 85719, USA

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了AI/ML在LSST暗能量科学合作中的应用机遇,强调了大规模贝叶斯推断、物理指导方法和主动学习等关键方法学优先事项,并讨论了新兴技术在重塑工作流程中的潜力。

Comments 84 pages. This is v1.0 of the DESC's white paper on AI/ML, a collaboration document that is being made public but which is not planned for submission to a journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00332 2026-01-22 cs.CL cs.AI 62%

Assertion-Conditioned Compliance: A Provenance-Aware Vulnerability in Multi-Turn Tool-Calling Agents

断言-条件合规:多轮工具调用代理中的一种意识-aware 的漏洞

Daud Waqas, Aaryamaan Golthi, Erika Hayashida, Huanzhi Mao

机构 * Monash University(墨尔本大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种新的评估方法A-CC,用于检测多轮工具调用代理在面对误导性断言时的鲁棒性问题,揭示了模型在用户和系统政策冲突下的脆弱性。

Comments 15 pages (incl. Appendix), 3 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15034 2026-01-22 cs.HC cs.AI 57%

Visual and Cognitive Demands of a Large Language Model-Powered In-vehicle Conversational Agent

基于大型语言模型的车载对话代理的视觉与认知需求

Chris Monk, Allegra Ayala, Christine S. P. Yu, Gregory M. Fitch, Dara Gruber

机构 * Exponent, Inc.(Exponent公司) Google, Inc.(Google公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 本研究评估了基于大型语言模型的车载对话代理在驾驶中的视觉与认知需求,发现其与免提通话在认知负荷上相当,且视觉需求较低,支持其在驾驶环境中的安全应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14827 2026-01-22 cs.AI 57%

Measuring and Aligning Abstraction in Vision-Language Models with Medical Taxonomies

利用医学分类法测量和对齐视觉-语言模型中的抽象能力

Ben Schaper, Maxime Di Folco, Bernhard Kainz, Julia A. Schnabel, Cosmin I. Bercea

机构 * 1 School of Computation, Information Technology, Technical University of Munich, Germany 2 Institute of Machine Learning in Biomedical Imaging, Helmholtz Munich, Germany 3 LTCI, Télécom Paris, Institut Polytechnique de Paris, France 4 Munich Center for Machine Learning (MCML) 5 School of Biomedical Engineering Imaging Sciences, King's College London, UK 6 Department of Artificial Intelligence in Biomedical Imaging, FAU Erlangen-Nuremberg, Germany 7 Department of Computing, Imperial College London, UK

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文提出通过医学分类法量化和缓解视觉-语言模型中的抽象错误,引入灾难性抽象错误概念,并通过风险约束阈值和分类法感知微调减少严重错误至2%以下。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03369 2026-01-22 cs.CV cs.CL 57%

RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models

RiskCueBench: 视频语言模型中早期风险信号的前瞻性推理基准测试

Sha Luo, Yogesh Prabhu, Timothy Ossowski, Kaiping Chen, Junjie Hu

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) University of California San Diego(加州大学圣地亚哥分校)

专题命中 安全评测 :safety(abstract);分类 cs.CL

AI总结 RiskCueBench通过标注早期风险信号片段,评估视频语言模型在预测未来风险事件中的能力,揭示了现有系统在解读动态情境方面的不足。

Comments *updated author email in this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11789 2026-01-22 cs.CL 57%

Personality Editing for Language Models through Adjusting Self-Referential Queries

通过调整自我参照查询进行语言模型的人格编辑

Seojin Hwang, Yumin Kim, Byeongjeong Kim, Donghoon Shin, Hwanhee Lee

机构 * Chung-Ang University(Chung-Ang 大学) University of Washington(华盛顿大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 PALETTE通过调整自我参照查询实现语言模型的人格编辑,仅需12个样本即可显著提升人格对齐性。

Comments Accepted to EACL 2026 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14460 2026-01-22 cs.IR 50%

Trust Me on This: A User Study of Trustworthiness for RAG Responses

相信我:对RAG响应可信度的用户研究

Weronika Łajewska, Krisztian Balog

专题命中 安全评测 :trustworthy(abstract)

AI总结 本研究通过用户实验探讨了不同解释类型如何影响用户对RAG响应的信任度,发现信任受响应清晰度和用户知识影响,而非单纯客观质量。

Comments This is the author's version of the work. The definitive version is published in: Proceedings of the 48th European Conference on Information Retrieval (ECIR '26), March 29-April 2, 2026, Delft, The Netherlands

详情

展开后加载摘要…

URL PDF HTML 收藏