arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-22 至 2026-01-22 共收录 40 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 2 篇

2502.18099 2026-01-22 cs.LG 79%

Stackelberg Self-Annotation: A Robust Approach to Data-Efficient LLM Alignment

Stackelberg 自注释:一种鲁棒的数据高效 LLM 对齐方法

Xu Chu, Zhixin Zhang, Tianyu Jia, Yujie Jin

机构 * Key Laboratory of High Confidence Software Technologies, Ministry of Education(高可信软件技术重点实验室,教育部) Center on Frontiers of Computing Studies, Peking University(计算前沿研究中心,北京大学) School of Computer Science, Peking University(计算机学院,北京大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

AI总结 本文提出Stackelberg自标注方法,通过少量人工标注和迭代自我标注,实现高效LLM对齐,减少对昂贵人工标注的依赖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05553 2026-01-22 cs.CL cs.CY 62%

Do Political Opinions Transfer Between Western Languages? An Analysis of Unaligned and Aligned Multilingual LLMs

政治观点在西方语言之间是否转移?对未对齐和对齐的多语言大语言模型的分析

Franziska Weeber, Tanise Ceron, Sebastian Padó

机构 * University of Stuttgart(斯图加特大学) Bocconi University(博科尼大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 研究探讨多语言大语言模型在五种西方语言中政治观点的转移情况,发现对齐前模型观点差异小,对齐后观点在各语言间均匀转移。

Comments EACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 9 篇

2507.09709 2026-01-22 cs.CL cs.LG 86%

Large Language Models Encode Semantics and Alignment in Linearly Separable Representations

大语言模型在线性可分的表示中编码语义和对齐

Baturay Saglam, Paul Kassianik, Blaine Nelson, Sajana Weerawardhena, Yaron Singer, Amin Karbasi

机构 * Yale University(耶鲁大学) Foundation AI – Cisco Systems Inc(Foundation AI – 卡西欧系统公司)

专题命中 安全训练 :alignment(title,abstract);safety(abstract);prompt injection(abstract);分类 cs.CL、cs.LG

AI总结 本研究发现大语言模型通过线性可分的表示编码语义和对齐,提出基于潜在空间的MLP探测器有效提升安全防护。

Comments IJCNLP and the Asian Chapter of ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14528 2026-01-22 cs.CR cs.LG 83%

LLM Security and Safety: Insights from Homotopy-Inspired Prompt Obfuscation

大语言模型的安全性与安全性:基于同调启发的提示混淆洞察

Luis Lazo, Hamed Jelodar, Roozbeh Razavi-Far

机构 * Canadian Institute for Cybersecurity(加拿大网络安全研究所) Faculty of Computer Science(计算机科学学院) University of New Brunswick(新 Brunswick大学)

专题命中 安全训练 :safety(title,abstract);trustworthy(abstract);分类 cs.LG

AI总结 本文提出基于同调启发的提示混淆框架,通过系统实验揭示LLM安全漏洞,强调对更强大防御机制和鲁棒性改进的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11689 2026-01-22 cs.CY cs.AI cs.CL 67%

Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot

面向社交与心理健康定制的生成式AI:一项现实世界试点

Thomas D. Hull, Lizhe Zhang, Patricia A. Arean, Matteo Malgaroli

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出了一种针对心理健康的生成式AI聊天机器人,通过现实世界试点研究证明其在心理健康支持方面的有效性与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17792 2026-01-22 cs.CL cs.AI cs.LG 67%

H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs

H3Fusion: 有助于、无害且诚实的对齐大语言模型融合

Selim Furkan Tekin, Fatih Ilhan, Tiansheng Huang, Sihao Hu, Yichang Xu, Zachary Yahn, Ling Liu

机构 * Georgia Institute of Technology, USA(佐治亚理工学院)

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 H3Fusion通过基于MoE的融合机制,在帮助性、无害性和诚实性方面优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10758 2026-01-22 cs.CY cs.AI 62%

Designing AI-Resilient Assessments Using Interconnected Problems: A Theoretically Grounded and Empirically Validated Framework

利用互联问题设计AI鲁棒评估:一种理论严谨且经验证实的框架

Kaihua Ding

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出了一种理论严谨且经验证实的框架,通过设计互联问题来提高评估的AI鲁棒性,挑战了开放性评估的鲁棒性假设,并提供了实证支持和实用设计方法。

Comments 8 pages, 3 figures and 3 tables, under submission to IEEE FIE

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14686 2026-01-22 cs.AI cs.LG 62%

IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization

IB-GRPO: 通过基于指标的群体相对策略优化对基于大语言模型的学习路径推荐进行对齐

Shuai Wang, Yaoming Yang, Bingdong Li, Hao Hao, Aimin Zhou

机构 * East China Normal University(东华大学)

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 IB-GRPO通过基于指标的群体相对策略优化,解决大语言模型在学习路径推荐中的多目标对齐问题,提升学习效果与教学目标的一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15115 2026-01-22 cs.CV 50%

Training-Free and Interpretable Hateful Video Detection via Multi-stage Adversarial Reasoning

无需训练的多阶段对抗推理 hateful 视频检测

Shuonan Yang, Yuchen Zhang, Zeyu Fu

机构 * Multimodal Intelligence Lab, Department of Computer Science, University of Exeter, United Kingdom(埃克塞特大学计算机科学系多模态智能实验室) Institute for Analytics and Data Science, University of Essex, United Kingdom(埃塞克斯大学分析与数据科学研究所)

专题命中 安全训练 :safety(abstract)

AI总结 MARS通过多阶段对抗推理框架实现无需训练的可解释仇恨视频检测,提升检测可靠性与透明度。

Comments Accepted at ICASSP 2026. \c{opyright} 2026 IEEE. This is the author accepted manuscript. The final published version will be available via IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14324 2026-01-22 cs.HC 50%

When Generative AI Is Intimate, Sexy, and Violent: Examining Not-Safe-For-Work (NSFW) Chatbots on FlowGPT

当生成AI变得亲密、性感和暴力:检验FlowGPT上的不适宜内容聊天机器人

Xian Li, Yuanning Han, Di Liu, Pengcheng An, Shuo Niu

专题命中 安全训练 :safety(abstract)

AI总结 研究分析FlowGPT上NSFW聊天机器人的类型及用户互动,揭示虚拟亲密、性幻想、暴力表达等特征,提出对聊天机器人设计和内容审核的启示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06136 2026-01-22 cs.CR 50%

"Abuse Risks are Often Inherent to Product Features": Exploring AI Vendors' Bug Bounty and Responsible Disclosure Policies

滥用风险往往与产品特性相关

Yangheran Piao, Jingjie Li, Daniel W. Woods

专题命中 安全训练 :alignment(abstract)

AI总结 本文研究了AI供应商在漏洞赏金和负责任披露政策中的风险与实践,发现36%的供应商无明确政策,且数据访问等漏洞最常被纳入范围,而劫持等漏洞则被排除。

Comments At USENIX Security Symposium 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 1 篇

2509.01631 2026-01-22 cs.AI 83%

Unraveling LLM Jailbreaks Through Safety Knowledge Neurons

通过安全知识神经元揭开大语言模型劫持的面纱

Chongwen Zhao, Yutong Ke, Kaizhu Huang

机构 * Duke Kunshan University(杜克昆山大学)

专题命中 越狱攻击 :safety(title,abstract);jailbreak(abstract);分类 cs.AI

AI总结 本文提出SafeTuning方法,通过调整安全知识神经元激活来提升大语言模型对劫持攻击的防御能力,实验显示其在多个模型上均有效

Comments EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 1 篇

2508.04339 2026-01-22 cs.AI 57%

Deliberative Reasoning Network: An Uncertainty-Driven Paradigm for Belief-Tracked Inference with Pretrained Language Models

辩证推理网络:一种基于不确定性的信念跟踪推理范式,用于预训练语言模型

Anran Xu, Jincheng Wang, Baigen Cai, Tao Wen

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

AI总结 DRN通过不确定性最小化范式提升预训练语言模型的逻辑推理能力,实现高准确率和强泛化性能。

Comments This submission represents an early exploratory draft and was uploaded prematurely. The authors have decided to withdraw it because the current version does not accurately reflect the intended scope and technical formulation of the work, and may be misleading if cited

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 11 篇

2601.14728 2026-01-22 eess.AS cs.AI cs.CL cs.LG cs.SD 82%

AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering

AQAScore: 通过音频问答评估文本到音频生成中的语义对齐

Chun-Yi Kuan, Kai-Wei Chang, Hung-yi Lee

机构 * National Taiwan University(国立台湾大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 AQAScore通过音频问答评估文本到音频生成的语义对齐,利用大型语言模型的推理能力,有效捕捉语义不一致性。

Comments Manuscript in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15267 2026-01-22 cs.CY cs.AI cs.CL 67%

Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions

评估大型语言模型在法律应用中的表现:挑战、方法与未来方向

Yiran Hu, Huanghai Liu, Chong Wang, Kunran Li, Tien-Hsuan Wu, Haitao Li, Xinran Xu, Siqing Huo, Weihang Su, Ning Zheng, Siyuan Zheng, Qingyao Ai, Yun Liu, Renjun Bian, Yiqun Liu, Charles L. A. Clarke, Weixing Shen, Ben Kao

机构 * Tsinghua University(清华大学) The University of Hong Kong(香港大学) University of Waterloo(多伦多大学) Shanghai Jiaotong University(上海交通大学) Peking University(北京大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文探讨了大型语言模型在法律应用中的评估挑战与方法,分析了现有评估框架的局限性,并提出了未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14620 2026-01-22 eess.AS cs.AI cs.LG cs.SD 62%

Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models

缩放歧义:通过音频-语言模型增强语音情感识别中的人工标注

Wenda Zhang, Hongyu Jin, Siyi Wang, Zhiqiang Wei, Ting Dang

机构 * University of Melbourne(墨尔本大学) Xi'an Jiaotong University(西安交通大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文通过音频-语言模型生成合成标注,提升语音情感识别中模糊情感的真实分布可靠性,实验表明合成标注在低歧义区域效果显著,但高歧义情况需进一步优化。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14269 2026-01-22 cs.CL cs.AI 62%

The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues

支持的缓慢漂移:多轮心理健康大语言模型对话中的边界失败

Youyou Cheng, Zhuangwei Kang, Kerry Jiang, Chenyu Sun, Qiyang Pan

机构 * University of Incarnate Word School of Osteopathic Medicine(incarnate Word 学校医学部) Independent Researcher(独立研究者) Mayo Clinic(梅奥诊所)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出多轮压力测试框架,揭示LLMs在长对话中因安慰和同理心尝试导致的安全边界逐步侵犯问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14235 2026-01-22 astro-ph.IM astro-ph.CO cs.AI cs.LG stat.ML 62%

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

人工智能/机器学习在Rubin LSST暗能量科学合作中的机遇

LSST Dark Energy Science Collaboration, Eric Aubourg, Camille Avestruz, Matthew R. Becker, Biswajit Biswas, Rahul Biswas, Boris Bolliet, Adam S. Bolton, Clecio R. Bom, Raphaël Bonnet-Guerrini, Alexandre Boucaud, Jean-Eric Campagne, Chihway Chang, Aleksandra Ćiprijanović, Johann Cohen-Tanugi, Michael W. Coughlin, John Franklin Crenshaw, Juan C. Cuevas-Tello, Juan de Vicente, Seth W. Digel, Steven Dillmann, Mariano Javier de León Dominguez Romero, Alex Drlica-Wagner, Sydney Erickson, Alexander T. Gagliano, Christos Georgiou, Aritra Ghosh, Matthew Grayling, Kirill A. Grishin, Alan Heavens, Lindsay R. House, Mustapha Ishak, Wassim Kabalan, Arun Kannawadi, François Lanusse, C. Danielle Leonard, Pierre-François Léget, Michelle Lochner, Yao-Yuan Mao, Peter Melchior, Grant Merz, Martin Millon, Anais Möller, Gautham Narayan, Yuuki Omori, Hiranya Peiris, Laurence Perreault-Levasseur, Andrés A. Plazas Malagón, Nesar Ramachandra, Benjamin Remy, Cécile Roucelle, Jaime Ruiz-Zapatero, Stefan Schuldt, Ignacio Sevilla-Noarbe, Ved G. Shah, Tjitske Starkenburg, Stephen Thorp, Laura Toribio San Cipriano, Tilman Tröster, Roberto Trotta, Padma Venkatraman, Amanda Wasserman, Tim White, Justine Zeghal, Tianqing Zhang, Yuanyuan Zhang

机构 * Université Paris Cité, CNRS, CEA, Astroparticule et Cosmologie, F-75013 Paris, France Department of Physics, University of Michigan, Ann Arbor, MI 48109, USA Leinweber Institute of Theoretical Physics, University of Michigan, Ann Arbor, MI 48109, USA Argonne National Laboratory, 9700 South Cass Avenue, Lemont, IL 60439, USA Cavendish Astrophysics, University of Cambridge, Madingley Road, Cambridge CB3 0HA, UK Kavli Institute for Cosmology, University of Cambridge, Madingley Road, Cambridge CB3 0HA, UK SLAC National Accelerator Laboratory, Menlo Park, CA 94025, USA Department of Computer Science, University of Milan, Milan, Italy Université Paris Cité, CNRS, Astroparticule et Cosmologie, F-75013 Paris, France Université Paris-Saclay, CNRS/IN2P3, IJCLab, 91405 Orsay, France Department of Astronomy Astrophysics, University of Chicago, Chicago, IL 60637, USA Kavli Institute for Cosmological Physics, University of Chicago, Chicago, IL 60637, USA NSF-Simons AI Institute for the Sky (SkAI), 172 E. Chestnut St., Chicago, IL 60611, USA Fermi National Accelerator Laboratory, P.O. Box 500, Batavia, IL 60510, USA Universit\'e Clermont-Auvergne, CNRS, LPCA, 63000 Clermont-Ferrand, France Kavli Institute for Particle Astrophysics Cosmology, Stanford University, Stanford, CA 94305, USA Department of Physics, Stanford University, 382 Via Pueblo Mall, Stanford, CA 94305, USA Engineering Faculty, Universidad Autonoma de San Luis Potosi, Zona Universitaria, San Luis Potosi, 78290, Mexico Stanford Artificial Intelligence Laboratory, Stanford University, Stanford, CA 94305, USA Kavli Institute of Cosmological Physics, University of Chicago, Chicago, IL 60637, USA The NSF AI Institute for Artificial Intelligence Center for Astrophysics Harvard \& Smithsonian, 60 Garden Street, Cambridge, MA 02138, USA Department of Physics Kavli Institute for Astrophysics Space Research, Massachusetts Institute of Technology, Cambridge, MA 02139, USA Institut de Física d'Altes Energies (IFAE), The Barcelona Institute of Science Institute of Astronomy Kavli Institute for Cosmology, University of Cambridge, Madingley Road, Cambridge, CB3 0HA, UK Imperial Centre for Inference Cosmology (ICIC), Imperial College London, Blackett Laboratory, Prince Consort Road, London SW7 2AZ, UK Data Science Institute, The University of Chicago, Chicago, IL 60615, USA Department of Physics, The University of Texas at Dallas, Richardson, TX 75080, USA Department of Physics, Duke University, Durham, NC 27708, USA Université Paris-Saclay, Université Paris Cité, CEA, CNRS, AIM, F-91191 Gif-sur-Yvette, France School of Mathematics, Statistics Physics, Newcastle University, Newcastle upon Tyne, NE1 7RU, United Kingdom Department of Astrophysical Sciences, Princeton University, Princeton, NJ 08544, USA Astronomy, University of the Western Cape, Bellville, Cape Town, 7535, South Africa Astronomy, University of Utah, Salt Lake City, UT 84112, USA Department of Astrophysical Sciences, Princeton University, Peyton Hall, Princeton, NJ 08544, USA Department of Astronomy, University of Illinois Urbana Champaign, 1002 W. Green St., Urbana, IL, 61801, USA Institute for Particle Physics Astrophysics, ETH Zürich, Wolfgang-Pauli-Strasse 27, CH-8093 Zurich, Switzerland Swinburne University of Technology, Hawthorn, Victoria 3122, Australia Ciela - Montr\'eal Institute for Astrophysical Data Analysis Mila - Quebec Artificial Intelligence Institute, Montréal, QC H2S 3H1, Canada Advanced Research Computing Centre, University College London, 90 High Holborn, London WC1V 6LJ, UK Finnish Centre for Astronomy with ESO (FINCA), University of Turku, FI-20014 Turku, Finland Department of Physics, P.O. Box 64, University of Helsinki, FI-00014 Helsinki, Finland Astronomy, Northwestern University, Evanston, IL, USA Center for Interdisciplinary Exploration Research in Astrophysics, Northwestern University, Evanston, IL, USA Scientific Data Science, International School for Advanced Study, Via Bonomea 265, I-34136 Trieste, Italy Department of Statistics, University of Michigan, Ann Arbor, MI 48109, USA PITT PACC, University of Pittsburgh, Pittsburgh, PA 15260, USA NSF NOIRLab, 950 N. Cherry Ave., Tucson, AZ 85719, USA

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了AI/ML在LSST暗能量科学合作中的应用机遇,强调了大规模贝叶斯推断、物理指导方法和主动学习等关键方法学优先事项,并讨论了新兴技术在重塑工作流程中的潜力。

Comments 84 pages. This is v1.0 of the DESC's white paper on AI/ML, a collaboration document that is being made public but which is not planned for submission to a journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00332 2026-01-22 cs.CL cs.AI 62%

Assertion-Conditioned Compliance: A Provenance-Aware Vulnerability in Multi-Turn Tool-Calling Agents

断言-条件合规:多轮工具调用代理中的一种意识-aware 的漏洞

Daud Waqas, Aaryamaan Golthi, Erika Hayashida, Huanzhi Mao

机构 * Monash University(墨尔本大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种新的评估方法A-CC,用于检测多轮工具调用代理在面对误导性断言时的鲁棒性问题,揭示了模型在用户和系统政策冲突下的脆弱性。

Comments 15 pages (incl. Appendix), 3 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15034 2026-01-22 cs.HC cs.AI 57%

Visual and Cognitive Demands of a Large Language Model-Powered In-vehicle Conversational Agent

基于大型语言模型的车载对话代理的视觉与认知需求

Chris Monk, Allegra Ayala, Christine S. P. Yu, Gregory M. Fitch, Dara Gruber

机构 * Exponent, Inc.(Exponent公司) Google, Inc.(Google公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 本研究评估了基于大型语言模型的车载对话代理在驾驶中的视觉与认知需求,发现其与免提通话在认知负荷上相当,且视觉需求较低,支持其在驾驶环境中的安全应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14827 2026-01-22 cs.AI 57%

Measuring and Aligning Abstraction in Vision-Language Models with Medical Taxonomies

利用医学分类法测量和对齐视觉-语言模型中的抽象能力

Ben Schaper, Maxime Di Folco, Bernhard Kainz, Julia A. Schnabel, Cosmin I. Bercea

机构 * 1 School of Computation, Information Technology, Technical University of Munich, Germany 2 Institute of Machine Learning in Biomedical Imaging, Helmholtz Munich, Germany 3 LTCI, Télécom Paris, Institut Polytechnique de Paris, France 4 Munich Center for Machine Learning (MCML) 5 School of Biomedical Engineering Imaging Sciences, King's College London, UK 6 Department of Artificial Intelligence in Biomedical Imaging, FAU Erlangen-Nuremberg, Germany 7 Department of Computing, Imperial College London, UK

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文提出通过医学分类法量化和缓解视觉-语言模型中的抽象错误,引入灾难性抽象错误概念,并通过风险约束阈值和分类法感知微调减少严重错误至2%以下。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03369 2026-01-22 cs.CV cs.CL 57%

RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models

RiskCueBench: 视频语言模型中早期风险信号的前瞻性推理基准测试

Sha Luo, Yogesh Prabhu, Timothy Ossowski, Kaiping Chen, Junjie Hu

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) University of California San Diego(加州大学圣地亚哥分校)

专题命中 安全评测 :safety(abstract);分类 cs.CL

AI总结 RiskCueBench通过标注早期风险信号片段,评估视频语言模型在预测未来风险事件中的能力,揭示了现有系统在解读动态情境方面的不足。

Comments *updated author email in this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11789 2026-01-22 cs.CL 57%

Personality Editing for Language Models through Adjusting Self-Referential Queries

通过调整自我参照查询进行语言模型的人格编辑

Seojin Hwang, Yumin Kim, Byeongjeong Kim, Donghoon Shin, Hwanhee Lee

机构 * Chung-Ang University(Chung-Ang 大学) University of Washington(华盛顿大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 PALETTE通过调整自我参照查询实现语言模型的人格编辑,仅需12个样本即可显著提升人格对齐性。

Comments Accepted to EACL 2026 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14460 2026-01-22 cs.IR 50%

Trust Me on This: A User Study of Trustworthiness for RAG Responses

相信我:对RAG响应可信度的用户研究

Weronika Łajewska, Krisztian Balog

专题命中 安全评测 :trustworthy(abstract)

AI总结 本研究通过用户实验探讨了不同解释类型如何影响用户对RAG响应的信任度,发现信任受响应清晰度和用户知识影响,而非单纯客观质量。

Comments This is the author's version of the work. The definitive version is published in: Proceedings of the 48th European Conference on Information Retrieval (ECIR '26), March 29-April 2, 2026, Delft, The Netherlands

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 2 篇

2601.14298 2026-01-22 cs.CR cs.AI cs.CY 81%

Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)

大型语言模型(LLM)的信任、安全与伦理发展和部署的防护措施

Anjanava Biswas, Wrick Talukdar

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

AI总结 本文提出了一种灵活自适应序列机制,结合信任和安全模块,用于实现大型语言模型开发和部署的安全防护。

Journal ref Journal Of Science & Technology, Vol. 4 No. 6 (2023), 55-82

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11893 2026-01-22 cs.CY cs.AI 62%

Beyond Automation: Rethinking Work, Creativity, and Governance in the Age of Generative AI

超越自动化:在生成式人工智能时代重新思考工作、创造力与治理

Haocheng Lin

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文探讨生成式AI对工作、创造力和治理的影响,提出包容性AI治理框架,强调UBI与技能发展等的结合以应对AI带来的挑战。

Comments Improved structure and clarity of the introduction and literature review; explicit articulation of the paper's contributions; refined the integration of AI across labour, UBI, and governance

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 14 篇

2601.13664 2026-01-22 cs.CV 78%

VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement

VIAFormer:用于高保真体素细化的体素-图像对齐变换器

Tiancheng Fang, Bowen Pan, Lingxi Chen, Jiangjing Lyu, Chengfei Lyu, Chaoyue Niu, Fan Wu

机构 * Shanghai Jiao Tong University(上海交通大学) Alibaba Group(阿里巴巴集团)

专题命中 其他安全 :alignment(title,abstract)

AI总结 VIAFormer通过体素-图像对齐变换器实现高保真体素细化,结合图像索引、校正流目标和混合流变换器,提升多视角条件下体素修复的精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15286 2026-01-22 cs.CV cs.AI cs.LG cs.RO 62%

Iterative Refinement Improves Compositional Image Generation

迭代细化改进组合图像生成

Shantanu Jaiswal, Mihir Prabhudesai, Nikash Bhardwaj, Zheyang Qin, Amir Zadeh, Chuan Li, Katerina Fragkiadaki, Deepak Pathak

机构 * Carnegie Mellon University(卡内基梅隆大学) Lambda AI

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种迭代细化策略,通过视觉-语言模型反馈提升组合图像生成质量,实验显示在多个基准上均优于并行采样方法。

Comments Project webpage: https://iterative-img-gen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07690 2026-01-22 cs.CL cs.AI 62%

LoSemB: Logic-Guided Semantic Bridging for Inductive Tool Retrieval

LoSemB:基于逻辑的语义桥接用于归纳工具检索

Luyao Zhuang, Qinggang Zhang, Huachi Zhou, Yujing Zhang, Xiao Huang

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 LoSemB通过逻辑引导的语义桥接框架,在无需重新训练的情况下实现高效归纳工具检索,缓解分布偏移和相似度检索的脆弱性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14791 2026-01-22 cs.CV cs.LG 57%

Synthetic Data Augmentation for Multi-Task Chinese Porcelain Classification: A Stable Diffusion Approach

多任务中国瓷器分类的合成数据增强:一种稳定扩散方法

Ziyao Ling, Silvia Mirri, Paola Salomoni, Giovanni Delnevo

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本研究利用稳定扩散与LoRA生成合成数据,提升多任务中国瓷器分类的性能,验证了合成数据在增强真实数据集中的有效性及任务特定的改进效果。

详情

展开后加载摘要…

URL PDF HTML 收藏