发表机构
Hitachi America, Ltd(美国日立公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对金融披露文本的细粒度不一致分类问题,在SBID-FD基准上对比多种模型,发现紧凑型监督编码器效率较高,证据定位是关键瓶颈,需同时提升证据提取与类别推理能力。
AI 中文摘要
金融披露文本包含数值声明、时间表述、实体引用、政策承诺和风险描述,这些内容可能以不同性质的方式产生冲突。检测冲突只是第一步:审核工作流还需要确定冲突类型,因为数值、时间、引用、事实和规范性不一致需要不同的证据和下游检查。我们将此问题作为细粒度不一致分类进行研究。使用SBID-FD的固定5940个实例快照(SBID-FD是一个合成金融披露基准,包含11个不一致标签和配对的参考证据片段),我们在统一评估协议下比较了冻结嵌入分类器、微调编码器、证据增强分类器、提示式大语言模型和LoRA适配的生成模型。一个微调的3亿参数编码器达到61.9%的准确率,而LoRA适配的Qwen3.5-9B模型为61.5%,GPT-5.4为61.3%。由于这些系统在架构、监督、训练目标和输入格式上存在差异,我们将此视为紧凑型监督编码器的实用效率结果,而非关于模型规模的可控结论。提供黄金标准证据片段可将微调编码器提升至65.3%,而自动预测的片段可恢复部分但不完整的增益,表明定位质量仍是瓶颈。类级别分析显示,引用不一致对定位质量尤其敏感,而事实和逻辑不一致即使提供相关证据仍难以处理。总体而言,黄金标准、干扰项和按类别的分析将定位错误与剩余的类型区分错误区分开来,表明进展需要更强的证据提取能力和对密切相关不一致类别的更好推理能力。
英文摘要
Financial disclosures may contain numerical, temporal, referential, factual, and policy inconsistencies that require different evidence and reasoning to diagnose. We study fine-grained inconsistency classification: given a passage known to contain a conflict, the goal is to identify its type among 11 categories. Using a fixed snapshot of the synthetic SBID-FD benchmark, we compare frozen and fine-tuned encoders, evidence-augmented classifiers, prompted large language models, and LoRA-adapted generative models under a shared evaluation protocol. Task-specific adaptation yields large improvements over frozen representations, and a fine-tuned 300M encoder performs competitively with substantially larger prompted and adapted models. We further study whether localizing the conflicting claims improves classification through matched predicted-span, reference-span, and distractor-span conditions. The results show that automatically extracted evidence provides additional signal but recovers only part of the benefit obtained from reference spans. Per-class and confusion analyses further reveal that some inconsistency types are especially sensitive to localization quality, whereas others remain difficult even when the relevant evidence is supplied. These findings identify evidence localization and fine-grained type discrimination as distinct challenges and show that compact supervised encoders are strong baselines for this task.