» Articles » PMID: 23962615

FmcsR: Mismatch Tolerant Maximum Common Substructure Searching in R

Overview
Journal Bioinformatics
Specialty Biology
Date 2013 Aug 22
PMID 23962615
Citations 27
Authors
Affiliations
Soon will be listed here.
Abstract

Motivation: The ability to accurately measure structural similarities among small molecules is important for many analysis routines in drug discovery and chemical genomics. Algorithms used for this purpose include fragment-based fingerprint and graph-based maximum common substructure (MCS) methods. MCS approaches provide one of the most accurate similarity measures. However, their rigid matching policies limit them to the identification of perfect MCSs. To eliminate this restriction, we introduce a new mismatch tolerant search method for identifying flexible MCSs (FMCSs) containing a user-definable number of atom and/or bond mismatches.

Results: The fmcsR package provides an R interface, with the time-consuming steps of the FMCS algorithm implemented in C++. It includes utilities for pairwise compound comparisons, structure similarity searching, clustering and visualization of MCSs. In comparison with an existing MCS tool, fmcsR shows better time performance over a wide range of compound sizes. When mismatching of atoms or bonds is turned on, the compute times increase as expected, and the resulting FMCSs are often substantially larger than their strict MCS counterparts. Based on extensive virtual screening (VS) tests, the flexible matching feature enhances the enrichment of active structures at the top of MCS-based similarity search results. With respect to overall and early enrichment performance, FMCS outperforms most of the seven other VS methods considered in these tests.

Availability: fmcsR is freely available for all common operating systems from the Bioconductor site (http://www.bioconductor.org/packages/devel/bioc/html/fmcsR.html).

Contact: thomas.girke@ucr.edu.

Supplementary Information: Supplementary data are available at Bioinformatics online.

Citing Articles

uafR: An R package that automates mass spectrometry data processing.

Stratton C, Thompson Y, Zio K, Morrison 3rd W, Murrell E PLoS One. 2024; 19(7):e0306202.

PMID: 38968199 PMC: 11226021. DOI: 10.1371/journal.pone.0306202.


Chemical species recognition in an adaptive radiation of Hawaiian spiders (Araneae: Tetragnathidae).

Adams S, Gurajapu A, Qiang A, Gerbaulet M, Schulz S, Tsutsui N Proc Biol Sci. 2024; 291(2020):20232340.

PMID: 38593845 PMC: 11003775. DOI: 10.1098/rspb.2023.2340.


Developing an AI-based prediction model for anaphylactic shock from injection drugs using Japanese real-world data and chemical structure-based analysis.

Enokiya T, Ozaki K Daru. 2024; 32(1):253-262.

PMID: 38580799 PMC: 11087410. DOI: 10.1007/s40199-024-00511-4.


Repurposing Drugs for Senotherapeutic Effect: Potential Senomorphic Effects of Female Synthetic Hormones.

Bramwell L, Frankum R, Harries L Cells. 2024; 13(6.

PMID: 38534362 PMC: 10969307. DOI: 10.3390/cells13060517.


Discovery of a structural class of antibiotics with explainable deep learning.

Wong F, Zheng E, Valeri J, Donghia N, Anahtar M, Omori S Nature. 2023; 626(7997):177-185.

PMID: 38123686 PMC: 10866013. DOI: 10.1038/s41586-023-06887-8.