发表机构
FAU Erlangen; Pattern Recognition Lab(埃尔朗根-纽伦堡大学; 模式识别实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出基于视觉语言模型(VLM)的流水线,从历史拍卖目录中自动提取结构化拍品级元数据,对比不同部署模式的模型性能,为文化遗产机构提供可行的自动化分析方案。
AI 中文摘要
对于溯源研究和艺术市场研究而言,拍卖目录是追踪特定物品时空轨迹的重要资源。尽管历史拍卖目录遵循既定的领域惯例,但其内部格式仍存在高度变异性,且目前因拍卖拍品缺乏机器可读表示,大规模分析受到限制。我们提出一种流水线,用于从19和20世纪大型历史拍卖与销售目录数据库German Sales中自动提取结构化拍品级元数据。我们使用手动标注的代表性目录页面测试集,在不同提示策略和约束解码框架下评估视觉语言模型(VLM)。为反映文化遗产机构面临的实际约束,包括预算、计算资源和数据隐私要求,我们对方法在不同部署模式下进行基准测试,范围从商业提供商到本地托管的量化模型。我们发现商业端点建立了性能上限,而机构网关提供了可行的隐私保护替代方案。本地部署仍可行,但生成过程中必须严格执行输出结构以保证有效的JSON格式。尽管仍需要不同程度的人在回路修正,但这项工作表明,基于VLM的流水线可成功解锁历史拍卖目录以进行大规模自动化分析。
英文摘要
For provenance research and art market studies, auction catalogs are an essential resource to trace specific objects over time and space. While historical auction catalogs follow established domain conventions, their internal formatting remains highly variable, and their large-scale analysis is currently restricted by the lack of machine-readable representations of the auction lots. We propose a pipeline to automatically extract structured lot-level metadata from German Sales, a large database of historical auction and sales catalogs from the 19th and 20th centuries. Using a manually annotated test set of representative catalog pages, we evaluate Vision-Language Models (VLMs) under varying prompt strategies and constrained decoding frameworks. To reflect the practical constraints faced by cultural heritage institutions, including budget, compute resources, and data privacy requirements, we benchmark the methods across different deployment modes ranging from commercial providers to locally hosted, quantized models. We find that commercial endpoints establish the performance ceiling, while institutional gateways offer a viable, privacy-preserving alternative. Local deployments remain feasible, but strictly require enforcing the output structure during generation to guarantee a valid JSON format. While varying degrees of human-in-the-loop correction are still necessary, this work demonstrates that a VLM-based pipeline can successfully unlock historical auction catalogs for large-scale automated analysis.
CommentsAccepted at the VISART Workshop (Computer Vision for Art Analysis), ECCV 2026. 19 pages, 6 figures, 5 tables. Supplementary material included as an appendix. Code, benchmark data, and prompt templates: https://github.com/mathiaszinnen/auction-lot-extraction