
Abstract Motivation Apart from meta-predictors, most of today's methods for residue–residue contact prediction are based entirely on Direct Coupling Analysis (DCA) of correlated mutations in multiple sequence alignments (MSAs). These methods are on average ∼40% correct for the 100 strongest predicted contacts in each protein. The end-user who works on a single protein of interest will not know if predictions are either much more or much less correct than 40%, which is especially a problem if contacts are predicted to steer experimental research on that protein. Results We designed a regression model that forecasts the accuracy of residue–residue contact prediction for individual proteins with an average error of 7 percentage points. Contacts were predicted with two DCA methods (gplmDCA and PSICOV). The models were built on parameters that describe the MSA, the predicted secondary structure, the predicted solvent accessibility and the contact prediction scores for the target protein. Results show that our models can be also applied to the meta-methods, which was tested on RaptorX. Availability and implementation All data and scripts are available from http://comprec-lin.iiar.pwr.edu.pl/dcaQ/. Supplementary information Supplementary data are available at Bioinformatics online.
Models, Molecular, Bioinformatics, Radboudumc 19: Nanomedicine RIMLS: Radboud Institute for Molecular Life Sciences, CMBI - Radboud University Medical Center, Computational Biology, Proteins, Protein Structure, Secondary, Data Accuracy, Sequence Analysis, Protein, Mutation, Algorithms, Software
Models, Molecular, Bioinformatics, Radboudumc 19: Nanomedicine RIMLS: Radboud Institute for Molecular Life Sciences, CMBI - Radboud University Medical Center, Computational Biology, Proteins, Protein Structure, Secondary, Data Accuracy, Sequence Analysis, Protein, Mutation, Algorithms, Software
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 4 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
