InsClaimBench:跨决策链的保险理赔裁定基准测试
InsClaimBench: Benchmarking Insurance Claim Adjudication Across the Decision Chain
浏览论文内容
中文总结 AI 辅助
本文提出InsClaimBench,一个覆盖汽车、财产和健康保险的端到端理赔裁定基准,通过3,780个案例和86,656个原子规则判断,揭示LLM在决策链中可靠性逐步下降,强调一致组合与传播的必要性。
中文摘要 AI 辅助
近期,面向推理的大型语言模型(LLMs)的进展推动了对它们执行专业决策任务能力的评估日益增多。保险理赔裁定就是这样一项任务,要求模型在结构化的决策过程中将案件证据、保险规则、中间判断和赔付计算联系起来。我们推出了InsClaimBench,一个用于评估跨决策链的保险理赔裁定的端到端基准。基于真实理赔材料和结构化保险规则,InsClaimBench包含375个案件家族中的3,780个案例,涵盖汽车、财产和健康保险,共计86,656个原子规则判断。它从原子规则到裁定模块再到赔付决策和金额,对每个理赔进行评估,并通过受控的事实变体测试所需更改是否在各层级间正确传播。对六个LLM的评估揭示了沿决策链可靠性的逐步下降。赔付决策准确率在74.23%至80.19%之间,而联合决策-金额准确率降至47.54%至73.15%。较强的局部性能也无法确保案件级正确性:原子规则准确率达到95.48%,而规则向量精确匹配峰值仅为36.90%,且最常见的模块错误不一定与最终决策失败关联最大。在事实变化下,这些不一致进一步演变为传播失败:模块更新不如规则更新可靠,正确的局部判断仍可能产生错误的赔付,而正确的赔付可能掩盖中间错误。这些结果表明,可靠的理赔裁定需要跨决策链的一致组合与传播。
英文摘要
Recent advances in reasoning-oriented large language models (LLMs) have motivated increasing evaluation of their ability to perform professional decision tasks. Insurance claim adjudication is one such task, requiring models to connect case evidence, insurance rules, intermediate judgments, and payout calculations across a structured decision process. We introduce InsClaimBench, an end-to-end benchmark for evaluating insurance claim adjudication across the decision chain. Grounded in real claim materials and structured insurance rules, InsClaimBench contains 3,780 cases in 375 case families across auto, property, and health insurance, comprising 86,656 atomic rule judgments. It evaluates each claim from atomic rules through adjudication modules to payout decisions and amounts, with controlled factual variants testing whether required changes are correctly propagated across levels. Evaluation of six LLMs reveals a progressive loss of reliability along the decision chain. Payout-decision accuracy ranges from 74.23--80.19%, while joint decision--amount accuracy drops to 47.54--73.15%. Strong local performance also fails to ensure case-level correctness: atomic-rule accuracy reaches 95.48%, whereas rule-vector exact match peaks at only 36.90%, and the most frequent module errors are not necessarily those most associated with final-decision failure. Under factual changes, these inconsistencies further become propagation failures: module updates are less reliable than rule updates, correct local judgments can still yield incorrect payouts, and correct payouts can conceal intermediate errors. These results show that reliable claim adjudication requires consistent composition and propagation across the decision chain.
发表机构
- Fudan University(复旦大学)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。