Using the spike protein feature to predict infection risk and monitor the evolutionary dynamic of coronavirus

25/03/2020View on timeline

Title: Using the spike protein feature to predict infection risk and monitor the evolutionary dynamic of coronavirus

Authors: Xiao-Li Qiang, Peng Xu, Gang Fang, Wen-Bin Liu and Zheng Kou

Abstract

Background: Coronavirus can cross the species barrier and infect humans with a severe respiratory syndrome. SARS-CoV-2 with potential origin of bat is still circulating in China. In this study, a prediction model is proposed to evaluate the infection risk of non-human-origin coronavirus for early warning.

Methods: The spike protein sequences of 2666 coronaviruses were collected from 2019 Novel Coronavirus Resource (2019nCoVR) Database of China National Genomics Data Center on Jan 29, 2020. A total of 507 humanorigin viruses were regarded as positive samples, whereas 2159 non-human-origin viruses were regarded as negative. To capture the key information of the spike protein, three feature encoding algorithms (amino acid composition, AAC; parallel correlation-based pseudo-amino-acid composition, PC-PseAAC and G-gap dipeptide composition, GGAP) were used to train 41 random forest models. The optimal feature with the best performance was identified by the multidimensional scaling method, which was used to explore the pattern of human coronavirus.

Results: The 10-fold cross-validation results showed that well performance was achieved with the use of the GGAP (g = 3) feature. The predictive model achieved the maximum ACC of 98.18% coupled with the Matthews correlation coefficient (MCC) of 0.9638. Seven clusters for human coronaviruses (229E, NL63, OC43, HKU1, MERS-CoV, SARS-CoV, and SARS-CoV-2) were found. The cluster for SARS-CoV-2 was very close to that for SARS-CoV, which suggests that both of viruses have the same human receptor (angiotensin converting enzyme II). The big gap in the distance curve suggests that the origin of SARS-CoV-2 is not clear and further surveillance in the field should be made continuously. The smooth distance curve for SARS-CoV suggests that its close relatives still exist in nature and public health is challenged as usual.

Conclusions: The optimal feature (GGAP, g = 3) performed well in terms of predicting infection risk and could be used to explore the evolutionary dynamic in a simple, fast and large-scale manner. The study may be beneficial for the surveillance of the genome mutation of coronavirus in the field.

Keywords: Coronavirus, Cross-species infection, Spike protein, Machine learning

Using the spike protein feature to predict infection risk and monitor the evolutionary dynamic of coronavirus, by Xiao-Li Qiang, Peng Xu, Gang Fang, Wen-Bin Liu and Zheng Kou.

Using the spike protein feature to predict infection risk and monitor the evolutionary dynamic of coronavirus, by Xiao-Li Qiang, Peng Xu, Gang Fang, Wen-Bin Liu and Zheng Kou. (PDF)

0 comments

Comment
No comments avaliable.

Author

Info

Published in 31/03/2020

Updated in 19/02/2021

All events in the topic Selected articles:


02/05/2018 • 07:00:00MERS transmission and risk factors: a systematic review
27/05/2019 • 07:00:00Coronavirus envelope protein: current knowledge
21/03/2020 • 10:00:00COVID-19, SARS and MERS: are they closely related?
23/03/2020 • 09:00:00Management of Critically Ill Adults With COVID-19
27/03/2020 • 18:00:00The many estimates of the COVID-19 case fatality rate
27/03/2020 • 13:00:00Possible method for the production of a Covid-19 vaccine