发表机构
Technical University of Munich; Massachusetts Institute of Technology; Leuphana University Lüneburg(慕尼黑工业大学; 麻省理工学院; 吕讷堡勒法纳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出声音保留/代表比率指标,审计LLM摘要员工反馈中的表征偏差,发现批评比赞美更易保留、低频担忧大量丢失,并贡献实地证据与声音保留卡片。
AI 中文摘要
组织越来越多地通过大语言模型(LLM)摘要将员工反馈传递给领导者,这是一个未经审计的环节,使已经表达的声音被沉默。我们引入了一个用于摘要中表征偏差的声音保留/代表比率指标,并将其应用于一家全球专业服务公司的双语(英语/德语)2586条自由文本回复语料库。首先,员工提供批评比提供赞美更可靠(保留赞美的情况是批评的82倍)。其次,在45份领导者摘要中,该流程按流行度而非情感进行过滤:批评得以幸存,但仅被提及一次的担忧在86%的情况下被丢弃,且简短和仅德语的内容在同一维度上丢失(主题保留率0.14对0.74;德语方向性)。在控制频率后,情感没有独立影响;损害是由流行度驱动的,而仅基于情感的审计会忽略这一点。有针对性的提示只能恢复被命名的主题。我们贡献了该指标、实地证据以及一份细化的声音保留卡片。
英文摘要
Organizations increasingly route employee feedback to leaders through large language model (LLM) summaries, an unaudited layer that silences already-spoken voice. We introduce a Voice Retention / Representation Ratio metric for representational bias in summarization and apply it to a bilingual (English/German) corpus of 2,586 free-text responses from a global professional service company. First, employees supply criticism more reliably than praise (withholding praise is 82 times more common). Second, across 45 leader-summaries the pipeline filters by popularity, not sentiment: criticism survives, yet a concern voiced once is dropped 86% of the time, with short and German-only content lost on the same axis (theme retention 0.14 vs 0.74; German directional). Controlling for frequency, sentiment has no independent effect; the harm is prevalence-driven, which sentiment-only audits miss. A targeted prompt recovers only named themes. We contribute the metric, field evidence, and a disaggregated voice-retention card.
Comments10 pages, 3 figures, 3 tables. Accepted at the 60th Hawaii International Conference on System Sciences (HICSS 2027)