arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

领域偏移下的可靠金融命名实体识别

Reliable Financial Named Entity Recognition Under Domain Shift: Confidence Estimation and Selective Prediction

Zihao Zheng, Baichuan Li, Junyi Yao, Jiayu Long

arXiv 2608.19558首次发表:更新:

发表机构

Washington University in St. Louis; Southern Methodist University(圣路易斯华盛顿大学; 南卫理公会大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对领域偏移下的金融命名实体识别,评估BERT等模型的置信度信号,提出先检测严重分布偏移再应用置信度门控的分阶段部署策略,提升了预测可靠性。

AI 中文摘要

金融AI系统通常在一种文本语体上训练信息提取器,并将其部署到 filings(文件)、新闻和用户生成内容中,而标准F1分数无法表明当输入分布发生变化时,哪些预测仍可安全实现自动化。我们针对金融命名实体识别(NER)的置信度估计与选择性预测开展研究,采用涵盖SEC filings(美国证券交易委员会文件)、金融新闻和通用主题社交媒体的三级压力测试作为极端域外条件。我们使用5种推理时置信度信号、3种训练种子及bootstrap区间评估BERT标记器与LoRA微调的Qwen2.5-0.5B/1.5B模型。置信度排名本身会在分布偏移下发生变化:整体输出概率是最强的域内错误检测器,但在域外会性能下降;而实体跨度概率与自一致性更具鲁棒性,且自一致性无需事后拟合即可实现更好校准。弃权(不执行)操作使域内最高置信度40%输入的句子错误率从34.3%降至2%以下,且在金融新闻上仍有用,但在极端社交媒体偏移下无法恢复出有用的大的干净子集。这些结果催生了一种分阶段部署策略,即先在上游检测严重分布偏移,再应用预测级置信度门控。

英文摘要

Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, and standard F1 scores do not indicate which predictions remain safe to automate when that input distribution changes. We study confidence estimation and selective prediction for financial named entity recognition (NER) on a three-tier stress test spanning SEC filings, financial news, and general-topic social media as an extreme out-of-domain condition, evaluating a BERT tagger and LoRA-tuned Qwen2.5-0.5B/1.5B models with five inference-time confidence signals, three training seeds, and bootstrap intervals. Confidence rankings themselves change under shift: whole-output probability is the strongest in-domain error detector but deteriorates out of domain, whereas entity-span probability and self-consistency are more robust; self-consistency is also better calibrated without post-hoc fitting. Abstention reduces sentence error from 34.3% to below 2% on the highest-confidence 40% of in-domain inputs and remains useful on financial news, but recovers no usefully large clean subset under the extreme social-media shift. These results motivate a staged deployment strategy that detects severe distribution shift upstream before applying prediction-level confidence gating.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑