信用卡、困惑、计算与后果:我们能从语言模型推理中发现什么?
Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
浏览论文内容
中文总结 AI 辅助
该研究构建了首个基于真实信用卡协议的数值推理金融素养基准CreditCardQA,评估发现程序思维(PoT)提示可提升模型推理性能,错误多源于金融规则误用等,边缘案例易影响财务脆弱群体。
中文摘要 AI 辅助
我们推出了CreditCardQA,这是首个基于真实信用卡协议构建的、用于数值推理的金融素养基准。该数据集包含1800个问题,其中包括反映消费者自然询问费用、利息和还款情况的第一人称变体。我们在思维链(CoT)和程序思维(PoT)提示下评估了一系列大型语言模型和推理模型。总体而言,PoT提示能带来持续的性能提升,尤其对基线推理能力较弱的模型效果显著,还能缩小开源与闭源系统之间的差距。通过错误分析,我们发现错误并非主要源于算术,而是源于金融规则应用不当、遗漏条件以及对合同条款的误解。我们进一步分析了问题难度,发现比较、条件逻辑和货币约束类问题尤其具有挑战性。我们还发现,错误常出现在如逾期付款罚金或小额余额等边缘案例中,这些案例更可能影响低收入或财务脆弱群体。
英文摘要
We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains 1,800 questions, including first-person variants that reflect how consumers naturally ask about fees, interest, and payments. We evaluate a range of large language and reasoning models under Chain-of-Thought (CoT) and Program-of-Thought (PoT) prompting. Overall, PoT yields consistent performance gains, particularly for models with weaker baseline reasoning, and narrows gaps between open- and closed-source systems. Through error analysis, we show that failures arise less from arithmetic and more from misapplied financial rules, missed conditions, and misunderstandings of contractual terms. We further analyze question difficulty and find that comparisons, conditional logic, and monetary constraints are especially challenging. We also find that errors often arise in edge cases such as late-payment penalties or small-balance scenarios that are more likely to affect lower-income or financially vulnerable individuals.
发表机构
- College of Computing(计算机学院)
- Scheller College of Business(舍勒商学院)
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。