arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FinLifeBench:基于纵向银行对话的详尽生活事件历史与财务状态重建

FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue

Hangyeul Lee, Juyoung Oh, Jaeyong Ko, Sunmin Kim, Jaeik Park, Hyunkyu Kim, Jungmin Son, Pilsung Kang

arXiv 2609.01198首次发表:更新:

发表机构

Seoul National University; KakaoBank(首尔大学; KakaoBank(韩国 Kakao 银行))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FinLifeBench基准针对银行对话的两项纵向重建任务,测试11个LLM的性能,发现其存在事件遗漏、财务状态更新错误等问题,两项任务性能弱相关。

AI 中文摘要

重复的银行交互场景要求智能助手维护完整、最新且可追溯的客户记录,因为生活变化会在日常请求中偶然出现。现有基准侧重于问答、有限片段或定向检索,而非详尽的纵向重建。我们推出FinLifeBench,该基准针对相同的累积对话评估两项任务:重建每个生活事件实例及其首次建立的会话,以及在连续检查点重建完整的34路径财务状态。该基准包含来自20条独立合成轨迹的6000轮八次韩国银行会话,具备针对24种事件类型和34种状态路径的确定性、详尽的黄金标准,以及共识质量保证。在全上下文条件下对11个大型语言模型(LLM)进行测试,事件锚定召回率从15个会话时的0.591降至300个会话时的0.445。错误主要由事件遗漏而非锚定定位不佳导致,而财务状态重建常将已被取代或可能过时的信息视为当前信息;最佳GCA@15达到0.470。两项重建任务的性能仅存在弱相关性。这些结果表明,模型能够定位已恢复事件的证据,但仍无法维护完整且时间上有效的纵向记录。

英文摘要

Repeated banking interactions require assistants to maintain complete, current, and traceable customer records as life changes emerge incidentally in routine requests. Existing benchmarks emphasize question answering, bounded episodes, or targeted recall rather than exhaustive longitudinal reconstruction. We introduce FinLifeBench, which evaluates two tasks over the same cumulative dialogue: reconstructing every life-event instance with its first-establishing session and reconstructing a complete 34-path financial state at consecutive checkpoints. The benchmark contains 6,000 eight-turn Korean banking sessions from 20 independent synthetic trajectories, with deterministic, exhaustive gold for 24 event types and 34 state paths and consensus quality assurance. Across eleven LLMs under a full-context condition, event-anchor recall falls from 0.591 at 15 sessions to 0.445 at 300. Errors are driven primarily by omitted events rather than poor anchor localization, while financial-state reconstruction frequently treats superseded or potentially outdated information as current; the best GCA@15 reaches 0.470. Performance on the two reconstruction tasks is only weakly associated. These results show that models can localize evidence for recovered events while still failing to maintain complete and temporally valid longitudinal records.

Comments9 pages, 3 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑