Genomic Sketching with Multiplicities and Locality-sensitive Hashing Using Dashing 2
Overview
Authors
Affiliations
A genomic sketch is a small, probabilistic representation of the set of k-mers in a sequencing data set. Sketches are building blocks for large-scale analyses that consider similarities between many pairs of sequences or sequence collections. Although existing tools can easily compare tens of thousands of genomes, data sets can reach millions of sequences and beyond. Popular tools also fail to consider k-mer multiplicities, making them less applicable in quantitative settings. Here, we describe a method called Dashing 2 that builds on the SetSketch data structure. SetSketch is related to HyperLogLog (HLL) but discards use of leading zero count in favor of a truncated logarithm of adjustable base. Unlike HLL, SetSketch can perform multiplicity-aware sketching when combined with the ProbMinHash method. Dashing 2 integrates locality-sensitive hashing to scale all-pairs comparisons to millions of sequences. It achieves superior similarity estimates for the Jaccard coefficient and average nucleotide identity compared with the original Dashing, but in much less time while using the same-sized sketch. Dashing 2 is a free, open source software.
EvANI benchmarking workflow for evolutionary distance estimation.
Majidian S, Hwang S, Zakeri M, Langmead B bioRxiv. 2025; .
PMID: 40027788 PMC: 11870633. DOI: 10.1101/2025.02.23.639716.
Fractional hitting sets for efficient multiset sketching.
Rouze T, Martayan I, Marchet C, Limasset A Algorithms Mol Biol. 2025; 20(1):1.
PMID: 39923117 PMC: 11807336. DOI: 10.1186/s13015-024-00268-0.
-mer approaches for biodiversity genomics.
Jenike K, Campos-Dominguez L, Bodde M, Cerca J, Hodson C, Schatz M Genome Res. 2025; 35(2):219-230.
PMID: 39890468 PMC: 11874746. DOI: 10.1101/gr.279452.124.
Mumemto: efficient maximal matching across pangenomes.
Shivakumar V, Langmead B bioRxiv. 2025; .
PMID: 39803467 PMC: 11722392. DOI: 10.1101/2025.01.05.631388.
Combining DNA and protein alignments to improve genome annotation with LiftOn.
Chao K, Heinz J, Hoh C, Mao A, Shumate A, Pertea M Genome Res. 2024; 35(2):311-325.
PMID: 39730188 PMC: 11874971. DOI: 10.1101/gr.279620.124.