Genomic Sketching with Multiplicities and Locality-sensitive Hashing Using Dashing 2

Overview

Journal Genome Res

Specialty Genetics

Date 2023 Jul 6

PMID 37414575

Authors

Daniel N Baker

Ben Langmead

Affiliations

Soon will be listed here.

Abstract

A genomic sketch is a small, probabilistic representation of the set of k-mers in a sequencing data set. Sketches are building blocks for large-scale analyses that consider similarities between many pairs of sequences or sequence collections. Although existing tools can easily compare tens of thousands of genomes, data sets can reach millions of sequences and beyond. Popular tools also fail to consider k-mer multiplicities, making them less applicable in quantitative settings. Here, we describe a method called Dashing 2 that builds on the SetSketch data structure. SetSketch is related to HyperLogLog (HLL) but discards use of leading zero count in favor of a truncated logarithm of adjustable base. Unlike HLL, SetSketch can perform multiplicity-aware sketching when combined with the ProbMinHash method. Dashing 2 integrates locality-sensitive hashing to scale all-pairs comparisons to millions of sequences. It achieves superior similarity estimates for the Jaccard coefficient and average nucleotide identity compared with the original Dashing, but in much less time while using the same-sized sketch. Dashing 2 is a free, open source software.

Citing Articles

EvANI benchmarking workflow for evolutionary distance estimation.

Majidian S, Hwang S, Zakeri M, Langmead B bioRxiv. 2025; .

PMID: 40027788 PMC: 11870633. DOI: 10.1101/2025.02.23.639716.

Fractional hitting sets for efficient multiset sketching.

Rouze T, Martayan I, Marchet C, Limasset A Algorithms Mol Biol. 2025; 20(1):1.

PMID: 39923117 PMC: 11807336. DOI: 10.1186/s13015-024-00268-0.

-mer approaches for biodiversity genomics.

Jenike K, Campos-Dominguez L, Bodde M, Cerca J, Hodson C, Schatz M Genome Res. 2025; 35(2):219-230.

PMID: 39890468 PMC: 11874746. DOI: 10.1101/gr.279452.124.

Mumemto: efficient maximal matching across pangenomes.

Shivakumar V, Langmead B bioRxiv. 2025; .

PMID: 39803467 PMC: 11722392. DOI: 10.1101/2025.01.05.631388.

Combining DNA and protein alignments to improve genome annotation with LiftOn.

Chao K, Heinz J, Hoh C, Mao A, Shumate A, Pertea M Genome Res. 2024; 35(2):311-325.

PMID: 39730188 PMC: 11874971. DOI: 10.1101/gr.279620.124.

References

Yu Y, Weber G . HyperMinHash: MinHash in LogLog space. IEEE Trans Knowl Data Eng. 2024; 34(1):328-339. PMC: 10824537. DOI: 10.1109/tkde.2020.2981311. View

Kent W, Zweig A, Barber G, Hinrichs A, Karolchik D . BigWig and BigBed: enabling browsing of large distributed datasets. Bioinformatics. 2010; 26(17):2204-7. PMC: 2922891. DOI: 10.1093/bioinformatics/btq351. View

Jain C, Rodriguez-R L, Phillippy A, Konstantinidis K, Aluru S . High throughput ANI analysis of 90K prokaryotic genomes reveals clear species boundaries. Nat Commun. 2018; 9(1):5114. PMC: 6269478. DOI: 10.1038/s41467-018-07641-9. View

Song L, Florea L, Langmead B . Lighter: fast and memory-efficient sequencing error correction without counting. Genome Biol. 2014; 15(11):509. PMC: 4248469. DOI: 10.1186/s13059-014-0509-9. View

Jain C, Rhie A, Zhang H, Chu C, Walenz B, Koren S . Weighted minimizer sampling improves long read mapping. Bioinformatics. 2020; 36(Suppl_1):i111-i118. PMC: 7355284. DOI: 10.1093/bioinformatics/btaa435. View

Dilthey A, Jain C, Koren S, Phillippy A . Strain-level metagenomic assignment and compositional estimation for long reads with MetaMaps. Nat Commun. 2019; 10(1):3066. PMC: 6624308. DOI: 10.1038/s41467-019-10934-2. View

Baker D, Langmead B . Dashing: fast and accurate genomic distances with HyperLogLog. Genome Biol. 2019; 20(1):265. PMC: 6892282. DOI: 10.1186/s13059-019-1875-0. View

Criscuolo A . On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference. F1000Res. 2020; 9:1309. PMC: 7713896. DOI: 10.12688/f1000research.26930.1. View

OLeary N, Wright M, Brister J, Ciufo S, Haddad D, McVeigh R . Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Res. 2015; 44(D1):D733-45. PMC: 4702849. DOI: 10.1093/nar/gkv1189. View

10.

Zhao X . BinDash, software for fast genome distance estimation on a typical personal laptop. Bioinformatics. 2018; 35(4):671-673. DOI: 10.1093/bioinformatics/bty651. View

11.

LaPierre N, Alser M, Eskin E, Koslicki D, Mangul S . Metalign: efficient alignment-based metagenomic profiling via containment min hash. Genome Biol. 2020; 21(1):242. PMC: 7488264. DOI: 10.1186/s13059-020-02159-0. View

12.

Wang J, Zhang T, Song J, Sebe N, Shen H . A Survey on Learning to Hash. IEEE Trans Pattern Anal Mach Intell. 2017; 40(4):769-790. DOI: 10.1109/TPAMI.2017.2699960. View

13.

Gostincar C . Towards Genomic Criteria for Delineating Fungal Species. J Fungi (Basel). 2020; 6(4). PMC: 7711752. DOI: 10.3390/jof6040246. View

14.

Marcais G, DeBlasio D, Pandey P, Kingsford C . Locality-sensitive hashing for the edit distance. Bioinformatics. 2019; 35(14):i127-i135. PMC: 6612865. DOI: 10.1093/bioinformatics/btz354. View

15.

Schraudolph N . A fast, compact approximation of the exponential function. Neural Comput. 1999; 11(4):853-62. DOI: 10.1162/089976699300016467. View

16.

Edgar R . Local homology recognition and distance measures in linear time using compressed amino acid alphabets. Nucleic Acids Res. 2004; 32(1):380-5. PMC: 373290. DOI: 10.1093/nar/gkh180. View

17.

Ondov B, Starrett G, Sappington A, Kostic A, Koren S, Buck C . Mash Screen: high-throughput sequence containment estimation for genome discovery. Genome Biol. 2019; 20(1):232. PMC: 6833257. DOI: 10.1186/s13059-019-1841-x. View

18.

Ondov B, Treangen T, Melsted P, Mallonee A, Bergman N, Koren S . Mash: fast genome and metagenome distance estimation using MinHash. Genome Biol. 2016; 17(1):132. PMC: 4915045. DOI: 10.1186/s13059-016-0997-x. View

19.

Shaw J, Yu Y . Fast and robust metagenomic sequence comparison through sparse chaining with skani. Nat Methods. 2023; 20(11):1661-1665. PMC: 10630134. DOI: 10.1038/s41592-023-02018-3. View

20.

Steinegger M, Soding J . Clustering huge protein sequence sets in linear time. Nat Commun. 2018; 9(1):2542. PMC: 6026198. DOI: 10.1038/s41467-018-04964-5. View