A Software Tool 'CroCo' Detects Pervasive Cross-species Contamination in Next Generation Sequencing Data

Overview

Journal BMC Biol

Publisher Biomed Central

Specialty Biology

Date 2018 Mar 7

PMID 29506533

Citations 42

Authors

Paul Simion

Khalid Belkhir

Clementine Francois

Julien Veyssier

Jochen C Rink

Michael Manuel

Herve Philippe

Maximilian J Telford

Affiliations

Soon will be listed here.

Abstract

Background: Multiple RNA samples are frequently processed together and often mixed before multiplex sequencing in the same sequencing run. While different samples can be separated post sequencing using sample barcodes, the possibility of cross contamination between biological samples from different species that have been processed or sequenced in parallel has the potential to be extremely deleterious for downstream analyses.

Results: We present CroCo, a software package for identifying and removing such cross contaminants from assembled transcriptomes. Using multiple, recently published sequence datasets, we show that cross contamination is consistently present at varying levels in real data. Using real and simulated data, we demonstrate that CroCo detects contaminants efficiently and correctly. Using a real example from a molecular phylogenetic dataset, we show that contaminants, if not eliminated, can have a decisive, deleterious impact on downstream comparative analyses.

Conclusions: Cross contamination is pervasive in new and published datasets and, if undetected, can have serious deleterious effects on downstream analyses. CroCo is a database-independent, multi-platform tool, designed for ease of use, that efficiently and accurately detects and removes cross contamination in assembled transcriptomes to avoid these problems. We suggest that the use of CroCo should become a standard cleaning step when processing multiple samples for transcriptome sequencing.

Citing Articles

Multiple Displacement Amplification Facilitates SMRT Sequencing of Microscopic Animals and the Genome of the Gastrotrich Lepidodermella squamata (Dujardin 1841).

Roberts N, Gilmore M, Struck T, Kocot K Genome Biol Evol. 2024; 16(12).

PMID: 39590608 PMC: 11660948. DOI: 10.1093/gbe/evae254.

A Phylogenomic Backbone for Acoelomorpha Inferred From Transcriptomic Data.

Abalde S, Jondelius U Syst Biol. 2024; 74(1):70-85.

PMID: 39451056 PMC: 11809588. DOI: 10.1093/sysbio/syae057.

Patterns of molecular evolution in a parthenogenic terrestrial isopod ().

Yarbrough E, Chandler C PeerJ. 2024; 12:e17780.

PMID: 39071119 PMC: 11276757. DOI: 10.7717/peerj.17780.

PhyloAln: A Convenient Reference-Based Tool to Align Sequences and High-Throughput Reads for Phylogeny and Evolution in the Omic Era.

Huang Y, Sun Y, Li H, Li H, Pang H Mol Biol Evol. 2024; 41(7).

PMID: 39041199 PMC: 11287380. DOI: 10.1093/molbev/msae150.

Dataset of PLA2 family identified from transcriptomic high-throughput sequencing of (Scorpionida: Buthidae) venom gland.

Salabi F, Jafari H Data Brief. 2024; 55:110629.

PMID: 39022691 PMC: 11253220. DOI: 10.1016/j.dib.2024.110629.

References

Finet C, Timme R, Delwiche C, Marletaz F . Multigene phylogeny of the green lineage reveals the origin and diversification of land plants. Curr Biol. 2010; 20(24):2217-22. DOI: 10.1016/j.cub.2010.11.035. View

Xie Y, Wu G, Tang J, Luo R, Patterson J, Liu S . SOAPdenovo-Trans: de novo transcriptome assembly with short RNA-Seq reads. Bioinformatics. 2014; 30(12):1660-6. DOI: 10.1093/bioinformatics/btu077. View

Brandl H, Moon H, Vila-Farre M, Liu S, Henry I, Rink J . PlanMine--a mineable resource of planarian biology and biodiversity. Nucleic Acids Res. 2015; 44(D1):D764-73. PMC: 4702831. DOI: 10.1093/nar/gkv1148. View

Bray N, Pimentel H, Melsted P, Pachter L . Near-optimal probabilistic RNA-seq quantification. Nat Biotechnol. 2016; 34(5):525-7. DOI: 10.1038/nbt.3519. View

Sievers F, Wilm A, Dineen D, Gibson T, Karplus K, Li W . Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega. Mol Syst Biol. 2011; 7:539. PMC: 3261699. DOI: 10.1038/msb.2011.75. View

Borner J, Burmester T . Parasite infection of public databases: a data mining approach to identify apicomplexan contaminations in animal genome and transcriptome assemblies. BMC Genomics. 2017; 18(1):100. PMC: 5244568. DOI: 10.1186/s12864-017-3504-1. View

Wagner G, Kin K, Lynch V . Measurement of mRNA abundance using RNA-seq data: RPKM measure is inconsistent among samples. Theory Biosci. 2012; 131(4):281-5. DOI: 10.1007/s12064-012-0162-3. View

Laurin-Lemay S, Brinkmann H, Philippe H . Origin of land plants revisited in the light of sequence contamination and missing data. Curr Biol. 2012; 22(15):R593-4. DOI: 10.1016/j.cub.2012.06.013. View

Li B, Dewey C . RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics. 2011; 12:323. PMC: 3163565. DOI: 10.1186/1471-2105-12-323. View

10.

Simion P, Philippe H, Baurain D, Jager M, Richter D, Di Franco A . A Large and Consistent Phylogenomic Dataset Supports Sponges as the Sister Group to All Other Animals. Curr Biol. 2017; 27(7):958-967. DOI: 10.1016/j.cub.2017.02.031. View

11.

Kumar S, Jones M, Koutsovoulos G, Clarke M, Blaxter M . Blobology: exploring raw genome data for contaminants, symbionts and parasites using taxon-annotated GC-coverage plots. Front Genet. 2013; 4:237. PMC: 3843372. DOI: 10.3389/fgene.2013.00237. View

12.

Merchant S, Wood D, Salzberg S . Unexpected cross-species contamination in genome sequencing projects. PeerJ. 2014; 2:e675. PMC: 4243333. DOI: 10.7717/peerj.675. View

13.

Whelan N, Kocot K, Moroz L, Halanych K . Error, signal, and the placement of Ctenophora sister to all other animals. Proc Natl Acad Sci U S A. 2015; 112(18):5773-8. PMC: 4426464. DOI: 10.1073/pnas.1503453112. View

14.

Laumer C, Bekkouche N, Kerbl A, Goetz F, Neves R, Sorensen M . Spiralian phylogeny informs the evolution of microscopic lineages. Curr Biol. 2015; 25(15):2000-6. DOI: 10.1016/j.cub.2015.06.068. View

15.

Shen X, Hittinger C, Rokas A . Contentious relationships in phylogenomic studies can be driven by a handful of genes. Nat Ecol Evol. 2017; 1(5):126. PMC: 5560076. DOI: 10.1038/s41559-017-0126. View

16.

Roure B, Rodriguez-Ezpeleta N, Philippe H . SCaFoS: a tool for selection, concatenation and fusion of sequences for phylogenomics. BMC Evol Biol. 2007; 7 Suppl 1:S2. PMC: 1796611. DOI: 10.1186/1471-2148-7-S1-S2. View

17.

Podar M, Haddock S, Sogin M, Harbison G . A molecular phylogenetic framework for the phylum Ctenophora using 18S rRNA genes. Mol Phylogenet Evol. 2001; 21(2):218-30. DOI: 10.1006/mpev.2001.1036. View

18.

Bolger A, Lohse M, Usadel B . Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014; 30(15):2114-20. PMC: 4103590. DOI: 10.1093/bioinformatics/btu170. View

19.

Struck T . The impact of paralogy on phylogenomic studies - a case study on annelid relationships. PLoS One. 2013; 8(5):e62892. PMC: 3647064. DOI: 10.1371/journal.pone.0062892. View

20.

Gouy M, Guindon S, Gascuel O . SeaView version 4: A multiplatform graphical user interface for sequence alignment and phylogenetic tree building. Mol Biol Evol. 2009; 27(2):221-4. DOI: 10.1093/molbev/msp259. View