发表机构
University of Malta(马耳他大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出vcf2db样本中心关系框架,通过模块化建模和共享样本标识符实现多组学数据集成检索,在选择性查询上优于现有工具。
AI 中文摘要
背景:高通量分子数据的快速增长要求系统能够跨组学层实现高效检索、集成和可扩展性。传统的基于文件的工作流程因碎片化存储和临时查询而阻碍了跨模态分析和可重复性。现有少数针对基因组变异数据的工具优先考虑模块化多组学集成。我们开发了vcf2db,一种以样本为中心的关系框架,将每个组学模态建模为以生物样本为中心的独特但可链接的组件。本研究评估了这种设计是否能在实现扩展到其他分子层的同时提供有竞争力的基因组检索。结果:我们使用1000基因组计划的欧洲子集(502个样本,2500万个变异)为注释的VCF数据实现了一个概念验证的基因组模式和摄取流程。我们在七项检索任务上将其与三种成熟的VCF导向工具进行了基准测试:坐标过滤、注释驱动查询、基因型提取和聚合。在受控条件下,vcf2db在选择性查询上表现强劲,通常在坐标和注释过滤器上优于其他系统,并且在基因型检索方面保持可用。聚合密集型任务效率较低,表明存在优化目标。我们还通过添加一个合成的转录组层验证了模块化可扩展性,无需修改基因组表,通过共享样本标识符链接各层。结论:vcf2db支持直接以锚定在共享样本标识符上的SQL查询形式进行跨层检索,从而实现了基于文件的方法难以实现的集成多组学访问。
英文摘要
Background: Rapid growth of high-throughput molecular data demands systems for efficient retrieval, integration, and scalability across omics layers. Traditional file-based workflows hinder cross-modal analysis and reproducibility because of fragmented storage and ad-hoc querying. Few existing tools for genomic variation data prioritize modular multi-omics integration. We developed vcf2db, a sample-centric relational framework modeling each omics modality as a distinct but linkable component centered on biological samples. This study evaluates whether this design delivers competitive genomic retrieval while enabling extension to additional molecular layers. Results: We implemented a proof-of-concept genomic schema and ingestion pipeline for annotated VCF data using the European subset of the 1000 Genomes Project (502 samples, 25 million variants). We benchmarked it against three established VCF-oriented tools on seven retrieval tasks: coordinate filtering, annotation-driven queries, genotype extraction, and aggregation. Under controlled conditions, vcf2db performed strongly on selective queries, often outperforming other systems for coordinate and annotation filters, and remained usable for genotype retrieval. Aggregation-heavy tasks were less efficient, indicating optimization targets. We also validated modular extensibility by adding a synthetic transcriptomic layer without modifying genomic tables, linking layers via shared sample identifiers. Conclusion: vcf2db supports cross-layer retrieval directly as SQL queries anchored on shared sample identifiers, enabling integrated multi-omics access that is difficult with file-based approaches.
CommentsVersion submitted to BioData Mining as Methodology paper