TREAT:评估等价数学表示下的形式知识获取能力
TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations
浏览论文内容
中文总结 AI 辅助
本研究提出基准TREAT,测试大型语言模型从保等价数学变换中识别定理身份的能力,发现模型表现不佳且存在多种失败模式,凸显形式知识表示鲁棒性的重要性。
中文摘要 AI 辅助
AI系统日益在灵活的输入表示与下游工具使用的形式对象之间运行,一项关键挑战是识别陌生表述何时对应已知的形式对象。我们通过定理识别研究这一挑战:给定对定理条件的保等价变换,模型必须恢复与标准表述关联的定理身份。我们引入TREAT,这是一个用于评估大型语言模型能否从保等价的公式级变换中恢复已知定理身份的基准。TREAT不改写定理文本,而是改变定理条件本身的数学形式,通过残差方程、见证语句、优化恒等式、集合关系、算子形式及证明中间表征来表达已知结果。我们从抓取的定理页面出发,筛选出具有可用数学表达式形式的条目,提取规范的定理条件,并生成带有记录假设和逆映射的变换变体。最终语料库包含737个定理身份和29480个变换行。在测试面板上,最佳模型仅在60.73%的案例中检索到正确的定理身份,其他系统则显现出不同的失败模式,包括弃权(不执行)、错误检测和格式错误的输出。这些结果表明,在等价的表示变化下,定理知识可能较为脆弱。因此,TREAT提供了一个受控测试平台,用于评估对形式知识的表示鲁棒性获取,其对需要稳定目标对象、显式等价关系、验证程序及可审计评分的领域具有更广泛的相关性。
英文摘要
AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is recognizing when an unfamiliar formulation denotes a known formal object. We study this challenge through theorem recognition: given an equivalence-preserving transformation of a theorem condition, a model must recover the theorem identity associated with the standard statement. We introduce TREAT, a benchmark for evaluating whether large language models can recover known theorem identities from equivalence-preserving formula-level transformations. Rather than paraphrasing theorem text, TREAT changes the mathematical form of theorem conditions themselves, expressing known results through residual equations, witness statements, optimization identities, set relations, operator forms, and proof-intermediate characterizations. Starting from scraped theorem pages, we filter for entries with usable mathematical expression forms, extract canonical theorem conditions, and generate transformed variants with recorded assumptions and inverse mappings. The final corpus contains 737 theorem identities and 29,480 transformed rows. On a test panel, the best model retrieves the correct theorem identity in only 60.73% of cases. Other systems reveal different failure modes, including abstention, wrong detection, and malformed outputs. These suggest that theorem knowledge can be fragile under equivalent changes in representation. TREAT therefore provides a controlled testbed for evaluating representation-robust access to formal knowledge, with broader relevance to domains that require stable target objects, explicit equivalence relations, validation procedures, and auditable scoring.