发表机构
California Polytechnic State University(加州州立理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对立法证词中的自我介绍检测问题,构建含154万条话语的数据集,采用多种分类器,提出BERT增强型XGBoost模型,将F1值提升至0.9782,优化了政府会议发言者识别任务。
AI 中文摘要
自我介绍在立法委员会证词中十分常见,成功检测此类自我介绍并提取发言者姓名,对政府会议场景下的发言者识别任务极具帮助。本文提出一种基于机器学习的立法委员会证词自我介绍检测流水线:构建了包含五个州立法会议、共154万条话语的训练数据集,采用名称匹配启发式方法生成自动标签,训练了决策树、随机森林和XGBoost三种分类器以检测自我介绍并提取发言者姓名;特征集结合了词袋、位置上下文、结构信号、介绍性短语指示符及话语上下文特征。三种分类器中,XGBoost表现最优,F1值达0.9747且总错误最少,加入微调BERT的概率特征后性能进一步提升。作为扩展,用微调BERT分类器对全部候选数据集打分,并将BERT概率输出作为特征加入,该BERT增强型XGBoost模型将F1值从0.9747提升至0.9782,总测试错误从241降至207。与决策树基线(F1值0.9323)相比,性能提升主要源于话语上下文特征和提升集成策略,BERT提供了适度的补充信号。对误报的分析显示,少数误报是源数据姓名不一致导致的真实自我介绍被错误标注,表明所测指标略低于真实性能。
英文摘要
Self-introductions are common in legislative committee testimonies. Successfully detecting them and extracting the speaker's name is enormously helpful in the task of speaker identification in the context of government meetings. In this paper, we present a pipeline for detection of self-introductions in legislative committee testimony using machine learning. We construct a training dataset from 1.54 million utterances spanning five state legislative sessions, apply a name-matching heuristic to generate automatic labels, and train three classifiers: a decision tree, random forest, and XGBoost to find self-introductions and extract the speaker's name. We construct a feature set combining bag-of-words, positional context, structural signals, introductory phrase indicators, and discourse context features. Among the three classifiers, XGBoost achieves the best performance with an F1 score of 0.9747 and the fewest total errors; adding fine-tuned BERT probability features improves this further. As an extension, we score the full candidate dataset with a fine-tuned BERT classifier and add BERT probability outputs as features. This BERT-augmented XGBoost model improves F1 from 0.9747 to 0.9782 and reduces total test errors from 241 to 207. The primary gain over the decision tree baseline (F1 0.9323) is driven by discourse context features and the boosting ensemble strategy; BERT provides a modest complementary signal. Analysis of false positives reveals that a minority are genuine self-introductions mislabeled due to name inconsistencies in the source data, indicating that measured metrics modestly understate true performance.
CommentsPresented at AAIRC-AI4 conference, Las Vegas, NV, USA August 2026 https://ai4.io/aairc-ai4/