Urgent.News

What's breaking now, across thousands of outlets.

Science

SpectroVQ: Noise-Aware Compression of Proteomics Data via Vector-Quantized Deep Learning improves MS/MS data storage and Peptide Identification

The amount of proteomics data generated has dramatically grown for the past decade due to the wider accessibility to mass spectrometers and technological advances. Current data storage and compression techniques largely treat mass spectra as meaningless series of numbers, wasting storage on useless noise and limiting the compression ratio. Here, we present SpectroVQ, a noise-aware…

Mass spectrometry generates vast amounts of proteomics data each year, yet existing compression techniques focus on blindly treating mass spectra as a collection of numbers, wasting storage capacity on irrelevant noise. SpectroVQ offers a novel solution by employing a noise-aware vector-quantized autoencoder to both compress and denoise peptide tandem mass spectra without relying on any prior annotation.

By leveraging deep learning to exploit peptide fragmentation patterns, SpectroVQ demonstrates the ability to retain useful signals in spectra from diverse peptide ions, even those in unseen datasets.

In comparison to mzMLb, SpectroVQ achieves a remarkable 3-fold increase in compression ratio, all while maintaining an average cosine similarity of over 0.9 and 90% agreement in peptide identifications. This compression method not only reduces storage requirements but also enhances data quality. Moreover, the study introduces a novel approach to further improve peptide identifications by up to 15% through ordinary library searching.

This is accomplished by leveraging the SpectroVQ's tunable denoising capability, effectively increasing the accuracy of peptide identification in proteomics data storage.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in Science

More from Friday 11 September →