发表机构
Edge Hill University; Ermetal Otomotiv ve Eşya San. Tic. A.Ş.(埃奇希尔大学; 埃尔梅塔尔汽车与家具工业贸易股份公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文在6GB显存的本地GPU上评估五个7B-8B开放权重LLM的土耳其语长文档问答能力,提出证据标注协议区分检索与推理失败,端到端准确率49%-75%,且无检索配置显著优于TF-IDF基线。
AI 中文摘要
大多数支持土耳其语的大型语言模型(LLM)都是使用通用基准进行评估,而非针对长篇幅、结构复杂的领域文档。本文在资源受限的本地部署环境下,评估了五个开放权重的7B-8B模型在土耳其语文档问答任务中的表现。主要基准包含100个系统验证的问题,这些问题源自一份109页的工业研发报告,评估协议还通过第二份112页的公共部门报告和独立构建的100个问题集进行了复现。所有模型均在配备6 GB显存的NVIDIA RTX 3050笔记本电脑GPU上本地运行,采用受控提示、解码和4位量化。主要方法论贡献是一个证据标注的评估协议,该协议无需额外模型调用即可将检索失败与下游模型推理失败区分开来。在主要基准上,端到端准确率在49%至75%之间。此外,还使用95% Wilson区间和精确配对McNemar检验比较了七种词法、稠密和混合检索配置;在任一文档上,均无配置显著优于字符TF-IDF基线。证据召回率在两份报告中的饱和情况不同,表明检索和有效上下文容量对某些文档可能是约束因素,而对其他文档则不然。这些结果表明,在部署开放权重LLM处理土耳其语领域文档时,必须分别评估模型选择、检索行为和硬件限制。
英文摘要
Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain documents. This paper evaluates five open-weight 7B-8B models for Turkish document question answering under a resource-constrained local deployment setting. The primary benchmark contains 100 systematically validated questions derived from a 109-page industrial R&D report, and the evaluation protocol is replicated using a second 112-page public-sector report and an independently constructed 100-question set. All models are evaluated locally on an NVIDIA RTX 3050 laptop GPU with 6 GB VRAM using controlled prompting, decoding, and 4-bit quantisation. The principal methodological contribution is an evidence-annotated evaluation protocol that separates retrieval failure from downstream model reasoning failure without requiring additional model calls. On the primary benchmark, end-to-end accuracy ranges from 49% to 75%. Seven lexical, dense, and hybrid retrieval configurations are additionally compared using 95% Wilson intervals and exact paired McNemar tests; none significantly outperforms the character TF-IDF baseline on either document. Evidence recall saturates differently across the two reports, showing that retrieval and effective context capacity can be binding constraints for some documents but not others. These results demonstrate that model selection, retrieval behaviour, and hardware limits must be evaluated separately when deploying open-weight LLMs for Turkish domain documents.
Comments6