arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24002cs.AI

FinInteract:在模糊金融问答中基准化澄清与意图整合

FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering

Xinyu Wang, Tung Sum Thomas Kwok, Zhenghan Tai, Guang Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

针对金融问答中问题模糊性被忽视的问题,提出FinInteract双语基准,通过五类模糊分类和双解读配对,揭示单一标准答案幻觉,并证明澄清与意图整合对智能体性能的关键影响。

中文摘要 AI 辅助

大型语言模型智能体越来越多地通过搜索监管文件来回答金融问题。这类问题往往具有欺骗性的不充分指定:Meta Platforms的“营业收入”合并口径为467.5亿美元,但Family of Apps部门的营业收入为628.7亿美元,且每种解读都能在文件中得到精确验证。一个称职的智能体应识别出这种模糊性并主动提问,而不是贸然采用一种看似合理但并非本意的解读。现有金融基准无法衡量这一点,因为每个问题只有一个标准答案,无法区分那些解决了模糊性的智能体与那些猜测常见解读的智能体,我们将这一盲点称为“单一标准答案幻觉”。我们发布了FinInteract,一个包含173个实例的双语(英语/中文)基准,每个问题都配有一个默认解读和一个预期解读,涵盖五类模糊性分类体系,并评估智能体是否能够引出正确的澄清并随后整合该澄清。将相同的输出按默认解读而非预期解读重新评分,会使GPT-4o的准确率膨胀3.1倍,证实了这一幻觉。除此之外,我们发现,一旦提供了解读,模型的回答准确率超过90%,但当模型必须自行引出解读时,准确率最多仅为28.9%;在一个对实体范围和指标定义具有充分统计功效的分类体系中,针对不同类别的表现不均衡,而在其他类别中则属于探索性结果;此外,以模糊性类别为条件,在推理和训练时均能改善解析效果。

英文摘要

Large language model agents increasingly answer financial questions by searching regulatory filings. Such questions are often deceptively under-specified: Meta Platforms' "operating income" is $46.75B consolidated but $62.87B for the Family of Apps segment, and each reading is exactly verifiable against the filing. A capable agent should recognize the ambiguity and ask, rather than commit to a plausible but unintended reading. Existing financial benchmarks cannot measure this, because one gold answer per question cannot separate agents that resolve the ambiguity from those that guess the common reading, a blind spot we call the single-gold illusion. We release FinInteract, a bilingual (English/Chinese) benchmark of 173 instances that pairs each question with a default and an intended interpretation across a five-category ambiguity taxonomy, and grades whether an agent elicits the right clarification and then integrates it. Re-grading identical outputs against the default rather than the intended reading inflates GPT-4o's accuracy by 3.1 times, confirming the illusion. Beyond it, we find that models answer above 90% once the interpretation is supplied but at most 28.9% when they must elicit it themselves, that targeting is uneven across a taxonomy well powered for entity scope and metric definition and exploratory elsewhere, and that conditioning on the ambiguity category improves resolution at both inference and training time.

发表机构

  • McGill University(麦吉尔大学)
  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
  • University of Toronto(多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

↑