arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

INS-ActBench:一个用于评估大语言模型专业精算能力的综合基准

INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models

Changyu Chen, Chenwei Lin, Xian Xu

arXiv 2607.24273首次发表:更新:

发表机构

Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

介绍INS-ActBench这一综合基准评估大语言模型专业精算能力,含12050个问答对及三个子集,通过实验揭示前沿大语言模型在不同方面的能力表现,为精算大语言模型开发提供可重复基础。

AI 中文摘要

大语言模型在金融推理方面展现出强大潜力,但现有基准通常在不同场景下评估领域知识、数值推理、长文本理解和工具使用。这限制了对需要可审计、基于上下文且可执行工具决策的现实专业工作流程的评估能力。我们引入INS-ActBench,一个用于评估大语言模型专业精算能力的综合基准。它包含来自16个精算协会的公开考试和样题中的12050个问答对,涵盖三个子集。实验表明前沿大语言模型在标准化知识方面表现强劲,但在案例推理、基于工具的工作流程和司法管辖区敏感实践方面较弱。INS-ActBench为开发可靠的专业精算大语言模型提供了可重复的基础。

英文摘要

Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reasoning, long-context understanding, and tool use in separate settings. This limits their ability to assess realistic professional workflows that require auditable, context-grounded, and tool-executable decisions. We introduce \textbf{INS-ActBench}, a comprehensive benchmark for evaluating professional actuarial capability in LLMs. INS-ActBench contains 12,050 Q\&A pairs from public exams and sample questions released by 16 actuarial associations. It covers three subsets: \textbf{INS-Act-Know} for standardized actuarial knowledge, \textbf{INS-Act-Case} for long-context insurance case reasoning, and \textbf{INS-Act-Practice} for spreadsheet and R-code tasks with verifiable numerical outputs. Experiments on nine representative LLMs and human actuarial experts reveal a clear capability boundary: frontier LLMs perform strongly on standardized knowledge, but remain much weaker in case reasoning, tool-based workflows, and jurisdiction-sensitive practice. INS-ActBench provides a reproducible foundation for developing actuarial LLMs toward reliable professional assistance. The code is available at https://github.com/FDU-INS/INS-ActBench.

Comments18 pages, including appendices; 11 figures and 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑