Skill matching at scale: freelancer-project alignment for efficient multilingual candidate retrieval
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI
Comments Accepted by ACL 2024
Journal ref https://aclanthology.org/2024.acl-long.514/
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI
Comments This article has been accepted in CAAI Transactions on Intelligence Technology! Article ID: CIT2_12370, Article DOI: 10.1049/cit2.12370
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG
Comments 12 pages, 6 figure sets
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG
Comments Accept at ICML 2024
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI
Comments Proceedings of ArabicNLP 2023
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI
Comments Accepted to CIKM 2023 (Full Paper)
专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG
Comments 7 pages, 7 figures
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI
Comments Annual Meeting of The ACL 2023: Main conference long paper
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG
Comments Accepted to ICML 2023 and CVPR4XAI workshop 2023
专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG
Comments 40 pages
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG
Comments Fixed typos
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG
Comments ICPR 2022
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG
专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG
Comments International Science and Innovation Congress 2019, pp. 643-655, 13 pages, 10 figures
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.CY
量化理论上的AI对齐保证:贝叶斯说服中的接收者效用界
机构 * Cornell University(康奈尔大学)
专题命中 其他安全 :alignment(title,comments);分类 cs.AI
AI总结 通过贝叶斯说服模型,研究AI发送者优化错位目标时,人类接收者仍能获得多少有用信息,证明接收者效用比不超过3/2,并给出紧性下界。
Comments 12 pages, EC 2026 Poster and EC 2026 Incentive-Based AI Alignment Workshop Poster
机构 * Department of Linguistics and Modern Languages, The Chinese University of Hong Kong(语言学与现代语言系,香港中文大学) ; Brain and Mind Institute, The Chinese University of Hong Kong(脑与心智研究所,香港中文大学)
专题命中 其他安全 :alignment(title,comments);分类 cs.CL
Comments Hanlin Wu, Xufeng Duan, and Zhenguang Cai. 2025. Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 135-143, Albuquerque, New Mexico, USA. Association for Computational Linguistics. https://aclanthology.org/2025.cmcl-1.18/
Journal ref In Proceedings of CMCL, pages 135-143, ACL (2025)
大语言模型通过一种独特的统一机制生成有害内容
机构 * Kempner Institute, Harvard University(哈佛大学肯普纳研究所) ; Princeton University(普林斯顿大学) ; Harvard University(哈佛大学) ; Cohere ; Technion—IIT(以色列理工学院)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究通过权重剪枝揭示大语言模型中有害生成的内部结构,发现有害内容生成依赖于一组通用且与良性能力不同的权重,表明对齐训练重塑了有害表示,解释了领域微调引发的广泛对齐偏差。
并非所有大语言模型的推理都能在思维链中体现
机构 * New York University(纽约大学) ; University of Maryland(马里兰大学) ; TogetherAI
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究探讨大语言模型输出令牌是否体现所有推理,发现前沿模型存在利用无关填充令牌提升合成推理任务性能的不可见推理现象,评估多个模型,揭示填充令牌益处因模型和令牌而异,还表明其能服务隐藏目标,且强化学习等方法无法使填充令牌益处在测试时持续。
CuMA: 通过人口统计感知的适配器混合使大语言模型与稀疏文化价值观对齐
机构 * Southeast University(东南大学) ; ByteDance Inc.(字节跳动公司) ; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用重点实验室(东南大学),中华人民共和国教育部,中国)
专题命中 其他安全 :alignment(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG
AI总结 提出CuMA框架,通过人口统计感知路由将冲突梯度分离到专家子空间,解决密集模型在多文化对齐中的均值崩溃问题,在WorldValuesBench等基准上取得最优性能。
Comments ACL 2026 Main
检查你的大语言模型的秘密词典!五行代码揭示你的大语言模型学到了什么(包括它不应该学到的)
机构 * Mgnite Inc.(Mgnite公司)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 通过对lm_head权重矩阵进行奇异值分解(仅需五行PyTorch代码且无需模型推理),直接从模型权重中揭示可解释的语义子空间,并发现模型训练数据组成和策展哲学。
对Llama3-8b-Instruct自生成文本识别能力的检查与控制
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本研究探讨了LLM是否能识别自身生成的文本,发现Llama3-8b-Instruct模型能够区分自身输出与人类输出,并通过残差流中的特定向量控制其行为和感知,揭示了模型自我归属的认知机制。
Comments 10 pages, 13 figs, 2 tables, accepted as conference paper to ICLR 2025
Journal ref The Thirteenth International Conference on Learning Representations (ICLR 2025)
计划是什么?LLMs中隐式规划的度量及其在押韵生成和问答中的应用
机构 * HPI / University of Potsdam(HPI/波茨坦大学) ; Utrecht University(乌特勒支大学) ; Google DeepMind(谷歌DeepMind)
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出简单方法评估LLM隐式规划,通过押韵生成和问答案例展示其可扩展性,发现隐式规划在1B参数模型中普遍存在,为AI安全提供新视角。
Comments 41 pages, 34 figures, Accepted at ICLR 2026, Code available at https://github.com/Jim-Maar/implicit-planning-in-llms
刻画模型内禀技能
机构 * Virginia Tech(弗吉尼亚理工学院)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出模型内禀技能的刻画方法,通过从序列激活中恢复紧凑正交基,实现行为变化轴的自组织,验证了在推理和安全对齐中的有效性,优于人类定义的技能。
Comments We argue that when the goal is to intervene on model behavior, skill characterization should be *model-native*: grounded in the model's own representations rather than imposed through external ontologies
DiaBlo: 对角块足以用于微调
机构 * University at Albany, SUNY(纽约州立大学阿尔巴尼分校) ; IBM T. J. Watson Research Center(IBM 汤普逊·杰·沃森研究中心)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 DiaBlo是一种仅更新模型权重矩阵对角块的参数高效微调方法,通过消除低秩矩阵乘积需求,实现稳定收敛和高效训练。
Comments Accepted by ICLR 2026
迭代部署提升大语言模型的规划能力
机构 * University of Oxford(牛津大学) ; AI Sequrity Company(AI安全公司) ; UFRGS(乌拉圭联邦大学)
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 通过迭代部署大语言模型,利用用户编纂的数据提升规划能力,展现隐含奖励函数的强化学习机制,具有AI安全和训练制度替代的双重意义。