发表机构
Meta; Autodesk; New Jersey Institute of Technology; Salesforce; University of California, Los Angeles; Amazon AGI(Meta; 欧特克公司; 新泽西理工学院; salesforce公司; 加州大学洛杉矶分校; 亚马逊AGI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究基于TrustNLP研讨会六年论文,分析NLP可信领域从可解释性到生成式系统控制的转变,明确各信任维度的发展趋势及与领域整体的关联性,提出结构性见解与研究方向。
AI 中文摘要
可信自然语言处理研讨会(TrustNLP)自2021年起与重要的ACL会议联合举办,历经六届,论文数量从8篇增至41篇,记录了该领域从静态模型的事后可解释性向生成式系统的机制理解与主动控制的转变。我们综合了所有144篇会议论文的见解,依据成熟框架(TrustLLM、DecodingTrust)确立的六个信任维度对其进行分类,观察到其与能力涌现存在共现关系。首批高影响力聊天模型的发布同时激活了所有信任维度,而后续模型代际则将重点转向真实性与安全对齐。分类研究的分析显示,真实性是增长最快的维度(2021-2022年不存在,到2025-2026年占论文的37%),公平性仍是最稳定的主题,可解释性则呈现U型轨迹:因事后方法失去相关性而下降,但在2026年通过机制可解释性再度兴起。同期与ACL、NAACL、EACL及EMNLP(约2000篇论文)的跨会场对比表明,TrustNLP的主题分布与领域平均水平高度吻合。我们确定了四个结构性见解,并为研究界总结了可操作的方向。
英文摘要
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). We observe co-occurrences with capability emergence. The release of the first high-impact chat models activated all trust dimensions simultaneously, while subsequent model generations shifted focus toward truthfulness and safety alignment. Analysis from the classification study reveals that truthfulness is the fastest-growing dimension (absent in 2021-2022, comprising 37% of papers by 2025-2026), fairness remains the most consistent theme, and explainability exhibits a U-shaped trajectory; declining as post-hoc methods lost relevance but resurging in 2026 through mechanistic interpretability. A cross-venue comparison with ACL, NAACL, EACL, and EMNLP (~2K papers) in the same period shows that TrustNLP's topical distribution closely follows the field average. We identify four structural insights and conclude with actionable directions for the research community.
Comments17 pages, 2 figures, 3 tables. Submitted to ACL ARR August 2026 cycle (EACL 2027)