» Articles » PMID: 36482318

EndHiC: Assemble Large Contigs into Chromosome-level Scaffolds Using the Hi-C Links from Contig Ends

Overview
Publisher Biomed Central
Specialty Biology
Date 2022 Dec 9
PMID 36482318
Authors
Affiliations
Soon will be listed here.
Abstract

Background: The application of PacBio HiFi and ultra-long ONT reads have enabled huge progress in the contig-level assembly, but it is still challenging to assemble large contigs into chromosomes with available Hi-C scaffolding tools, which count Hi-C links between contigs using the whole or a large part of contig regions. As the Hi-C links of two adjacent contigs concentrate only at the neighbor ends of the contigs, larger contig size will reduce the power to differentiate adjacent (signal) and non-adjacent (noise) contig linkages, leading to a higher rate of mis-assembly.

Results: We design and develop a novel Hi-C based scaffolding tool EndHiC, which is suitable to assemble large contigs into chromosomal-level scaffolds. The core idea behind EndHiC, which distinguishes it from other Hi-C scaffolding tools, is using Hi-C links only from the most effective regions of contig ends. By this way, the signal neighbor contig linkages and noise non-neighbor contig linkages are separated more clearly. Benefiting from the increased signal to noise ratio, the reciprocal best requirement, as well as the robustness evaluation, EndHiC achieves higher accuracy for scaffolding large contigs compared to existing tools. EndHiC has been successfully applied in the Hi-C scaffolding of simulated data from human, rice and Arabidopsis, and real data from human, great burdock, water spinach, chicory, endive, yacon, and Ipomoea cairica, suggesting that EndHiC can be applied to a broad range of plant and animal genomes.

Conclusions: EndHiC is a novel Hi-C scaffolding tool, which is suitable for scaffolding of contig assemblies with contig N50 size near or over 10 Mb and N90 size near or over 1 Mb. EndHiC is efficient both in time and memory, and it is interface-friendly to the users. As more genome projects have been launched and the contig continuity constantly improved, we believe EndHiC has the potential to make a great contribution to the genomics field and liberate the scientists from labor-intensive manual curation works.

Citing Articles

A chromosome-scale reference assembly of Vigna radiata enables delineation of centromeres and telomeres.

Oraon P, Ambreen H, Yadav P, Ramarao S, Goel S Sci Data. 2025; 12(1):305.

PMID: 39979386 PMC: 11842788. DOI: 10.1038/s41597-025-04436-8.


A chromosome-scale reference genome of the Banna miniature inbred pig.

Chen H, Xu K, Yan C, Zhao H, Jiao D, Si S Sci Data. 2024; 11(1):1345.

PMID: 39695204 PMC: 11655879. DOI: 10.1038/s41597-024-04201-3.


A deep learning-based method enables the automatic and accurate assembly of chromosome-level genomes.

Jiang Z, Peng Z, Wei Z, Sun J, Luo Y, Bie L Nucleic Acids Res. 2024; 52(19):e92.

PMID: 39287126 PMC: 11514472. DOI: 10.1093/nar/gkae789.


Genome architecture of the allotetraploid wild grass Aegilops ventricosa reveals its evolutionary history and contributions to wheat improvement.

Liu Z, Yang F, Wan H, Deng C, Hu W, Fan X Plant Commun. 2024; 6(1):101131.

PMID: 39257004 PMC: 11783901. DOI: 10.1016/j.xplc.2024.101131.


The genomes of 5 underutilized Papilionoideae crops provide insights into root nodulation and disease resistance.

Yuan L, Lei L, Jiang F, Wang A, Chen R, Wang H Gigascience. 2024; 13.

PMID: 39190925 PMC: 11348429. DOI: 10.1093/gigascience/giae063.


References
1.
Lieberman-Aiden E, van Berkum N, Williams L, Imakaev M, Ragoczy T, Telling A . Comprehensive mapping of long-range interactions reveals folding principles of the human genome. Science. 2009; 326(5950):289-93. PMC: 2858594. DOI: 10.1126/science.1181369. View

2.
Dudchenko O, Batra S, Omer A, Nyquist S, Hoeger M, Durand N . De novo assembly of the genome using Hi-C yields chromosome-length scaffolds. Science. 2017; 356(6333):92-95. PMC: 5635820. DOI: 10.1126/science.aal3327. View

3.
Fan W, Wang S, Wang H, Wang A, Jiang F, Liu H . The genomes of chicory, endive, great burdock and yacon provide insights into Asteraceae palaeo-polyploidization history and plant inulin production. Mol Ecol Resour. 2022; 22(8):3124-3140. DOI: 10.1111/1755-0998.13675. View

4.
Zhou C, McCarthy S, Durbin R . YaHS: yet another Hi-C scaffolding tool. Bioinformatics. 2022; 39(1). PMC: 9848053. DOI: 10.1093/bioinformatics/btac808. View

5.
Ruan J, Li H . Fast and accurate long-read assembly with wtdbg2. Nat Methods. 2019; 17(2):155-158. PMC: 7004874. DOI: 10.1038/s41592-019-0669-3. View