Predicting the Predicted: A Comparison of Machine Learning-Based Collision Cross-Section Prediction Models for Small Molecules.

Sara M. de Cripan; Trisha Arora; Adrià Olomí; Núria Canela-Canela; Gary Siuzdak; Xavier Domingo-Almenara

Dades identificatives

Identificador: imarina:9368765

Handle: https://hdl.handle.net/20.500.11797/imarina9368765

Autors:
Sara M. de CripanTrisha AroraAdrià OlomíNúria Canela-CanelaGary SiuzdakXavier Domingo-Almenara

Resum:
The application of machine learning (ML) to -omics research is growing at an exponential rate owing to the increasing availability of large amounts of data for model training. Specifically, in metabolomics, ML has enabled the prediction of tandem mass spectrometry and retention time data. More recently, due to the advent of ion mobility, new ML models have been introduced for collision cross-section (CCS) prediction, but those have been trained with different and relatively small data sets covering a few thousands of small molecules, which hampers their systematic comparison. Here, we compared four existing ML-based CCS prediction models and their capacity to predict CCS values using the recently introduced METLIN-CCS data set. We also compared them with simple linear models and with ML models that used fingerprints as regressors. We analyzed the role of structural diversity of the data on which the ML models are trained with and explored the practical application of these models for metabolite annotation using CCS values. Results showed a limited capability of the existing models to achieve the necessary accuracy to be adopted for routine metabolomics analysis. We showed that for a particular molecule, this accuracy could only be improved when models were trained with a large number of structurally similar counterparts. Therefore, we suggest that current annotation capabilities will only be significantly altered with models trained with heterogeneous data sets composed of large homogeneous hubs of structurally similar molecules to those being predicted.
Altres:

Autor segons l'article: Sara M. de Cripan; Trisha Arora; Adrià Olomí; Núria Canela-Canela; Gary Siuzdak; Xavier Domingo-Almenara
Departament: Enginyeria Electrònica, Elèctrica i Automàtica
Autor/s de la URV: Arora, Trisha / Domingo Almenara, Xavier
Resum: The application of machine learning (ML) to -omics research is growing at an exponential rate owing to the increasing availability of large amounts of data for model training. Specifically, in metabolomics, ML has enabled the prediction of tandem mass spectrometry and retention time data. More recently, due to the advent of ion mobility, new ML models have been introduced for collision cross-section (CCS) prediction, but those have been trained with different and relatively small data sets covering a few thousands of small molecules, which hampers their systematic comparison. Here, we compared four existing ML-based CCS prediction models and their capacity to predict CCS values using the recently introduced METLIN-CCS data set. We also compared them with simple linear models and with ML models that used fingerprints as regressors. We analyzed the role of structural diversity of the data on which the ML models are trained with and explored the practical application of these models for metabolite annotation using CCS values. Results showed a limited capability of the existing models to achieve the necessary accuracy to be adopted for routine metabolomics analysis. We showed that for a particular molecule, this accuracy could only be improved when models were trained with a large number of structurally similar counterparts. Therefore, we suggest that current annotation capabilities will only be significantly altered with models trained with heterogeneous data sets composed of large homogeneous hubs of structurally similar molecules to those being predicted.
Àrees temàtiques: Analytical chemistry Astronomia / física Biodiversidade Biotecnología Chemistry, analytical Ciência da computação Ciência de alimentos Ciências agrárias i Ciências ambientais Ciências biológicas i Ciências biológicas ii Ciências biológicas iii Enfermagem Engenharias ii Engenharias iii Engenharias iv Ensino Farmacia General medicine Geociências Interdisciplinar Materiais Medicina i Medicina ii Química
Accès a la llicència d'ús: https://creativecommons.org/licenses/by/3.0/es/
Adreça de correu electrònic de l'autor: trisha.arora@estudiants.urv.cat xavier.domingo@urv.cat
Data d'alta del registre: 2024-11-23
Versió de l'article dipositat: info:eu-repo/semantics/publishedVersion
Referència a l'article segons font original: Analytical Chemistry. 96 (22): 9088-9096
Referència de l'ítem segons les normes APA: Sara M. de Cripan; Trisha Arora; Adrià Olomí; Núria Canela-Canela; Gary Siuzdak; Xavier Domingo-Almenara (2024). Predicting the Predicted: A Comparison of Machine Learning-Based Collision Cross-Section Prediction Models for Small Molecules.. Analytical Chemistry, 96(22), 9088-9096. DOI: 10.1021/acs.analchem.4c00630
URL Document de llicència: https://repositori.urv.cat/ca/proteccio-de-dades/
Entitat: Universitat Rovira i Virgili
Any de publicació de la revista: 2024
Tipus de publicació: Journal Publications

Paraules clau:

Analytical Chemistry,Chemistry, Analytical
Analytical chemistry
Astronomia / física
Biodiversidade
Biotecnología
Chemistry, analytical
Ciência da computação
Ciência de alimentos
Ciências agrárias i
Ciências ambientais
Ciências biológicas i
Ciências biológicas ii
Ciências biológicas iii
Enfermagem
Engenharias ii
Engenharias iii
Engenharias iv
Ensino
Farmacia
General medicine
Geociências
Interdisciplinar
Materiais
Medicina i
Medicina ii
Química
Documents:

DocumentPrincipal
Cerca a google

Repositori URV

Articles producció científica> Enginyeria Electrònica, Elèctrica i Automàtica

Predicting the Predicted: A Comparison of Machine Learning-Based Collision Cross-Section Prediction Models for Small Molecules.

Dades identificatives

Altres:

Paraules clau:

Documents:

Cerca a google