PipeMEM: A Framework to Speed Up BWA-MEM in Spark with Low Overhead

Overview

Journal Genes (Basel)

Publisher MDPI

Date 2019 Nov 7

PMID 31689965

Citations 6

Authors

Lingqi Zhang

Cheng Liu

Shoubin Dong

Affiliations

Soon will be listed here.

Abstract

(1) Background: DNA sequence alignment process is an essential step in genome analysis. BWA-MEM has been a prevalent single-node tool in genome alignment because of its high speed and accuracy. The exponentially generated genome data requiring a multi-node solution to handle large volumes of data currently remains a challenge. Spark is a ubiquitous big data platform that has been exploited to assist genome alignment in handling this challenge. Nonetheless, existing works that utilize Spark to optimize BWA-MEM suffer from higher overhead. (2) Methods: In this paper, we presented PipeMEM, a framework to accelerate BWA-MEM with lower overhead with the help of the pipe operation in Spark. We additionally proposed to use a pipeline structure and in-memory-computation to accelerate PipeMEM. (3) Results: Our experiments showed that, on paired-end alignment tasks, our framework had low overhead. In a multi-node environment, our framework, on average, was 2.27× faster compared with BWASpark (an alignment tool in Genome Analysis Toolkit (GATK)), and 2.33× faster compared with SparkBWA. (4) Conclusions: PipeMEM could accelerate BWA-MEM in the Spark environment with high performance and low overhead.

Citing Articles

Bioinformatics characterization of variants of uncertain significance in pediatric sensorineural hearing loss.

Clay S, Evans A, Zambrano R, Otohinoyi D, Hicks C, Tsien F Front Pediatr. 2024; 12:1299341.

PMID: 38450295 PMC: 10915201. DOI: 10.3389/fped.2024.1299341.

Multi-Omics Characterization of Circular RNA-Encoded Novel Proteins Associated With Bladder Outlet Obstruction.

Zhu B, Kang Z, Zhu S, Zhang Y, Lai X, Zhou L Front Cell Dev Biol. 2022; 9:772534.

PMID: 35071227 PMC: 8777291. DOI: 10.3389/fcell.2021.772534.

CircRNA expression profiling of PBMCs from patients with hepatocellular carcinoma by RNA-sequencing.

Han Z, Feng W, Hu R, Ge Q, Sun X, Ma W Exp Ther Med. 2021; 22(6):1467.

PMID: 34737807 PMC: 8561760. DOI: 10.3892/etm.2021.10902.

VC@Scale: Scalable and high-performance variant calling on cluster environments.

Ahmad T, Al Ars Z, Hofstee H Gigascience. 2021; 10(9).

PMID: 34494101 PMC: 8424057. DOI: 10.1093/gigascience/giab057.

Bioinformatics Accelerates the Major Tetrad: A Real Boost for the Pharmaceutical Industry.

Behl T, Kaur I, Sehgal A, Singh S, Bhatia S, Al-Harrasi A Int J Mol Sci. 2021; 22(12).

PMID: 34201152 PMC: 8227524. DOI: 10.3390/ijms22126184.

References

Liu Y, Schmidt B, Maskell D . CUSHAW: a CUDA compatible short read aligner to large genomes based on the Burrows-Wheeler transform. Bioinformatics. 2012; 28(14):1830-7. DOI: 10.1093/bioinformatics/bts276. View

Vurture G, Sedlazeck F, Nattestad M, Underwood C, Fang H, Gurtowski J . GenomeScope: fast reference-free genome profiling from short reads. Bioinformatics. 2017; 33(14):2202-2204. PMC: 5870704. DOI: 10.1093/bioinformatics/btx153. View

Li H, Durbin R . Fast and accurate long-read alignment with Burrows-Wheeler transform. Bioinformatics. 2010; 26(5):589-95. PMC: 2828108. DOI: 10.1093/bioinformatics/btp698. View

Wiewiorka M, Messina A, Pacholewska A, Maffioletti S, Gawrysiak P, Okoniewski M . SparkSeq: fast, scalable and cloud-ready tool for the interactive genomic data analysis with nucleotide precision. Bioinformatics. 2014; 30(18):2652-3. DOI: 10.1093/bioinformatics/btu343. View

Jourdren L, Bernard M, Dillies M, Le Crom S . Eoulsan: a cloud computing-based framework facilitating high throughput sequencing analyses. Bioinformatics. 2012; 28(11):1542-3. DOI: 10.1093/bioinformatics/bts165. View

Pireddu L, Leo S, Zanetti G . SEAL: a distributed short read mapping and duplicate removal tool. Bioinformatics. 2011; 27(15):2159-60. PMC: 3137215. DOI: 10.1093/bioinformatics/btr325. View

Torri F, Dinov I, Zamanyan A, Hobel S, Genco A, Petrosyan P . Next generation sequence analysis and computational genomics using graphical pipeline workflows. Genes (Basel). 2012; 3(3):545-75. PMC: 3490498. DOI: 10.3390/genes3030545. View

Feuerriegel S, Schleusener V, Beckert P, Kohl T, Miotto P, Cirillo D . PhyResSE: a Web Tool Delineating Mycobacterium tuberculosis Antibiotic Resistance and Lineage from Whole-Genome Sequencing Data. J Clin Microbiol. 2015; 53(6):1908-14. PMC: 4432036. DOI: 10.1128/JCM.00025-15. View

Abuin J, Pichel J, Pena T, Amigo J . SparkBWA: Speeding Up the Alignment of High-Throughput DNA Sequencing Data. PLoS One. 2016; 11(5):e0155461. PMC: 4868289. DOI: 10.1371/journal.pone.0155461. View

10.

Abuin J, Pichel J, Pena T, Amigo J . BigBWA: approaching the Burrows-Wheeler aligner to Big Data technologies. Bioinformatics. 2015; 31(24):4003-5. DOI: 10.1093/bioinformatics/btv506. View

11.

Langmead B, Salzberg S . Fast gapped-read alignment with Bowtie 2. Nat Methods. 2012; 9(4):357-9. PMC: 3322381. DOI: 10.1038/nmeth.1923. View

12.

Smith T, Waterman M . Identification of common molecular subsequences. J Mol Biol. 1981; 147(1):195-7. DOI: 10.1016/0022-2836(81)90087-5. View

13.

Nordberg H, Bhatia K, Wang K, Wang Z . BioPig: a Hadoop-based analytic toolkit for large-scale sequence data. Bioinformatics. 2013; 29(23):3014-9. DOI: 10.1093/bioinformatics/btt528. View

14.

Chiang C, Layer R, Faust G, Lindberg M, Rose D, Garrison E . SpeedSeq: ultra-fast personal genome analysis and interpretation. Nat Methods. 2015; 12(10):966-8. PMC: 4589466. DOI: 10.1038/nmeth.3505. View

15.

Simonyan V, Mazumder R . High-Performance Integrated Virtual Environment (HIVE) Tools and Applications for Big Data Analysis. Genes (Basel). 2014; 5(4):957-81. PMC: 4276921. DOI: 10.3390/genes5040957. View

16.

Weese D, Holtgrewe M, Reinert K . RazerS 3: faster, fully sensitive read mapping. Bioinformatics. 2012; 28(20):2592-9. DOI: 10.1093/bioinformatics/bts505. View

17.

Altschul S, Gish W, Miller W, Myers E, Lipman D . Basic local alignment search tool. J Mol Biol. 1990; 215(3):403-10. DOI: 10.1016/S0022-2836(05)80360-2. View

18.

Zhao M, Lee W, Garrison E, Marth G . SSW library: an SIMD Smith-Waterman C/C++ library for use in genomic applications. PLoS One. 2013; 8(12):e82138. PMC: 3852983. DOI: 10.1371/journal.pone.0082138. View