arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GAGR-Lab:评估联合空间几何与分析函数推理

GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning

Jingyao Zhang, Yun Li, Lu Han

arXiv 2610.10201首次发表:更新:

AI 中文总结

GAGR-Lab是一个评估联合空间几何与分析函数推理能力的框架,通过游戏场景和Rust轨迹执行进行测试,试点显示模型无目标命中,但分析搜索控制成功,贡献了可操作的研究框架和前瞻性计划。

AI 中文摘要

联合空间几何与分析函数推理需要将感知到的空间配置转换为符号函数,该函数执行后的曲线需满足几何约束。我们提出了GAGR-Lab,一个通过笛卡尔游戏场景、显式函数语义和权威的Rust轨迹执行来衡量这种能力的框架。它区分了空间感知、度量基础、几何关系、函数解释、函数构建和约束综合。我们指定了四个可配置的场景难度预设和一个前瞻性的24单元诊断设计,同时仅报告实际评估的子集。对一种托管模型(Llama 3.2 11B Vision Instruct)进行的有界试点,使用两个API凭据作为执行副本,产生了72个平衡游戏,共432次尝试,429个有效提供者响应,且没有目标命中;探索性的普通函数提示变体也未能命中,而结构化定位接口未产生可评分的输出。一个特权分析搜索控制独立地在来自300个生成场景的600个方向性案例上取得成功,具有精确的可重复性和1,200次成功的垂直反射或平移检查。该框架分离了服务可靠性、符号合规性和几何成功,并保留了精确的模型可见输入和实现路径。一个分阶段的协议概述了诊断校准、保留复制、多模型比较和配对鲁棒性测试。贡献是一个可操作的研究框架,附带已执行的试点和明确标识的前瞻性研究计划;完整的难度矩阵和比较性模型结果仍未测试。

英文摘要

Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and authoritative Rust trajectory execution. It distinguishes spatial perception, metric grounding, geometric relations, function interpretation, function construction, and constrained synthesis. We specify four configurable scene-difficulty presets and a prospective 24-cell diagnostic design, while reporting only the subset actually evaluated. A bounded pilot of one hosted model (Llama 3.2 11B Vision Instruct) using two API credentials as execution replicas yields 72 balanced games with 432 attempts, 429 valid provider responses, and no target hits; exploratory ordinary-function prompt variants also fail to hit, while the structured localization interface yields no scoreable outputs. A privileged analytic search control independently succeeds on 600 directional cases from 300 generated scenes, with exact repeatability and 1,200 successful vertical-reflection or translation checks. The framework separates serving reliability, symbolic compliance, and geometric success, and preserves exact model-visible inputs and realized paths. A staged protocol outlines diagnostic calibration, held-out replication, multi-model comparison, and paired robustness tests. The contribution is an operational research framework with an executed pilot and a clearly identified prospective study plan; the full difficulty matrix and comparative model results remain untested.

Comments15 pages, 1 figure, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑