Integrating molecular fingerprints with machine learning for accurate neurotoxicity prediction: an observational study

Gao, Yilin1; Mu, Jingyuan2; Liu, Kaifeng1,*; Wang, Min3,4,*


1School of Life Sciences, Jilin University, Changchun, Jilin Province, China

2Department of Computer Science & Engineering, Texas A&M University, TX, USA

3International Research Center for Nano Handling and Manufacturing of China, Changchun University of Science and Technology, Changchun, Jilin Province, China

4Ministry of Education Key Laboratory for Cross-Scale Micro and Nano Manufacturing, Changchun University of Science and Technology, Changchun, Jilin Province, China


*Correspondence to: Kaifeng Liu, PhD, liukf1220@mails.jlu.edu.cn; Min Wang, PhD, wangm@cust.edu.cn.


Funding:The Program of Science and Technology Development Plan of Jilin Province, No. 20210204009YY (to MW) and the Program of Department of Education of Jilin Province, No. JJKH 20210853KJ (to MW).


This is an open access journal, and articles are distributed under the terms of the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 License (http://creativecommons.org/licenses/by-nc-sa/4.0/), which allows others to remix, tweak, and build upon the work non-commercially, as long as appropriate credit is given and the new creations are licensed under the identical terms.


Advanced Technology in Neuroscience 2(3):p 109-115, September 2025. | DOI: 10.4103/ATN.ATN-D-24-00034


Abstract

Neurotoxicity refers to harmful changes in the structure and function of the central and/or peripheral nervous system caused by exposure to chemical, physical, or biological factors. These changes can manifest as organic damage, functional disorders, and behavioral alterations. Traditional biological testing methods for neurotoxicity are often expensive, labor-intensive, and time-consuming. Additionally, existing predictive models rely on small datasets and lack comprehensive, large-scale, and professional algorithms and platforms for accurately predicting brain toxicity. In this study, we integrate molecular fingerprints and descriptors—namely, MorganFP, MACCS, RDKFP, and TopologicalTorsionFP—with machine learning models such as Random Forest and Support Vector Machine, as well as deep learning models such as Graph Neural Networks, to develop a neurotoxicity prediction model. The optimal model, MorganFP-SVM, demonstrates excellent performance in predicting neurotoxicity, achieving an accuracy of 86.56%. This significantly outperforms other neurotoxicity prediction models, including ADMETlab 3.0. When compared to existing neurotoxicity prediction models DINeuroT and ADMETlab 3.0, the MorganFP-SVM model exhibits superior performance across multiple key evaluation metrics, offering greater balance and comprehensiveness as a reliable tool for neurotoxicity prediction. The DINeuroT model shows weaknesses in specificity and Matthews correlation coefficient, while ADMETlab 3.0, though strong in sensitivity, has limited capacity to accurately identify negative samples. Overall, the comprehensive advantages of the MorganFP-SVM model in neurotoxicity prediction establish it as an essential tool in the field. It not only provides higher predictive accuracy but also demonstrates strong stability and reliability across various evaluation metrics. This model offers a promising approach for assessing neurotoxicity risks related to drug development and environmental pollutants, thereby providing a scientific basis for public health decision-making.


中文摘要

神经毒性是指因接触化学、物理或生物因素而导致中枢和/或周围神经系统的结构和功能发生有害变化,包括机体损伤、功能紊乱和行为变化。传统的神经毒性生物检测方法成本高昂、耗费大量人力和时间。此外,现有的预测模型都是基于小型数据集,缺乏全面、大规模、专业化的脑毒性预测算法和平台。这项研究采用分子指纹和描述符[包括 MorganFP、MACCS、RDKFP 和 TopologicalTorsionFP]与机器学习模型[如随机森林(RF)和支持向量机(SVM))和深度学习模型(如图神经网络(GNN))相结合的新技术,开发了一种神经毒性预测模型。最优模型 MorganFP-SVM 在预测神经毒性方面表现出色,准确率高达 86.56%,明显优于 admetlab3.0 等其他神经毒性预测模型。通过与现有神经毒性预测模型DINeuroT和ADMETlab3.0的对比,Morgan-SVM模型在多个关键评价指标上表现出更优的性能,具有更强的平衡性和综合性,是一个可靠的神经毒性预测工具。DINeuroT模型在特异性和马修斯相关系数(MCC)上表现较弱,而ADMETlab3.0虽然在敏感性上表现良好,但特异性地识别阴性样本的能力有限。总体而言,Morgan-SVM模型在神经毒性预测上的全面优势,使其成为该领域的重要工具。它不仅提供了更高的预测准确性,而且在多个评价指标上显示出强大的稳定性和可靠性。该模型为药物开发和环境污染物神经毒性风险的评估提供了有前途的方法,并为公共卫生相关决策提供了科学依据。