arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

平衡RAG流水线中的推理与硬件约束:面向乌克兰多领域文档理解

Balancing Reasoning and Hardware Constraints in RAG Pipelines for Ukrainian Multi-Domain Document Understanding

Illya Havrylov

arXiv 2609.22124首次发表:更新:

发表机构

National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute”; Educational and Scientific Institute for Applied System Analysis (IASA)(乌克兰国立技术大学“伊戈尔·西科斯基基辅理工学院”; 应用系统分析教育与科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对乌克兰多领域文档理解任务,提出资源高效的混合RAG流水线,利用BM25、BGE-M3和交叉编码器,采用4位量化LapaLLM 12B模型,在硬件限制下平衡推理与稳定性,获得私有分数0.8095,排名第10。

AI 中文摘要

本文描述了提交至UNLP 2026多领域文档理解共享任务(Shared Task)的系统。该挑战要求在严格的9小时离线Kaggle执行限制内,从多样化的乌克兰语PDF文档语料库中提取精确答案、文档ID和页码。在隐藏私有测试集评估期间,扫描文档的光学字符识别(OCR)成为严重瓶颈,由于顺序单线程执行,消耗了总时间预算中的5-7小时。这一开销严格限制了剩余用于大型语言模型(LLM)推理的时间,仅剩约两小时来处理500个问题。为确保流水线无超时完成,我们开发了一种资源高效的混合检索增强生成(RAG)流水线,利用BM25、BGE-M3和交叉编码器重排序。我们没有部署参数繁重的推理模型(如DeepSeek R1),因为这些模型持续超时,而是通过此http URL在双NVIDIA T4 GPU上使用了4位量化的LapaLLM 12B模型。优先考虑流水线稳定性而非多步推理,我们的系统取得了0.8095的私有分数,在15个活跃团队中排名第10。

英文摘要

This paper describes the system submitted to the UNLP 2026 Shared Task on Multi-Domain Document Understanding. The challenge required extracting precise answers, document IDs, and page numbers from a diverse corpus of Ukrainian PDF documents within a strict 9-hour offline Kaggle execution limit. During evaluation on the hidden private test set, optical character recognition (OCR) of scanned documents emerged as a severe bottleneck, consuming 5-7 hours of the total time budget due to sequential single-threaded execution. This overhead strictly limited the remaining time for Large Language Model (LLM) inference to approximately two hours for 500 questions. To guarantee pipeline completion without timeouts, we developed a resource-efficient Hybrid Retrieval-Augmented Generation (RAG) pipeline utilizing BM25, BGE-M3, and Cross-Encoder reranking. Rather than deploying parameter-heavy reasoning models (e.g., DeepSeek R1) which consistently timed out, we utilized a 4-bit quantized LapaLLM 12B model via llama.cpp on dual NVIDIA T4 GPUs. Prioritizing pipeline stability over multi-step reasoning, our system achieved a Private Score of 0.8095, placing 10th out of 15 active teams.

Comments5 pages, 2 tables. Technical report based on UNLP 2026 Shared Task submission

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑