Benchmarking Graph Neural Networks for Multi-Omics Cancer Subtyping using Methylation and Gene Expression Profiles
Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal…
Biomedical researchers have shown growing interest in employing Graph Neural Networks (GNNs) to extract deeper insights into molecular processes and disease mechanisms. Despite this growing interest, there has been a lack of comprehensive evaluation of GNNs for graph signal classification in the biomedical domain. To address this gap, a team of scientists conducted a benchmarking study, comparing multiple GNN architectures on a Protein-Protein Interaction (PPI) network to predict subtypes of Kidney Renal Clear Cell Carcinoma and Breast Cancer.
The researchers evaluated several GNN models, incorporating diverse data modalities and graph structures, including skip connections. The findings revealed that none of the GNNs outperformed the structure-agnostic Multi-Layer Perceptron baseline, but all models were capable of handling bimodal data, which consists of gene methylation and expression profiles. Furthermore, the GNNs were able to provide explainability through the PPI network.
The study offers practical guidelines for applying GNNs to graph signal processing tasks in the context of cancer classification. Depending on the specific dataset and PPI structure used, different data modalities demonstrated varying performance. Overall, the researchers recommend the use of ChebNet, which consistently outperformed both Graph Convolutional Network and Graph Attention Network in cancer subtype prediction.
Additionally, they suggest employing GNN architectures with a simple flattening readout layer, as these models tend to offer better classification accuracy and faster training times compared to those utilizing global average pooling. Lastly, the researchers found that residual connections had minimal impact on the classification performance.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.