Lossless compression of protein databases for efficient and accurate metagenomic sequence classification with Centrifuger
We present a lossless compression algorithm for indexing a protein database while supporting fast taxonomic classification in the method Centrifuger. The algorithm is a new scheme of the previously proposed run-block compression algorithm to reduce the size of the Ferragina-Manzini (FM) index, and it scales better with alphabet size than the original. On the RefSeq prokaryotic and viral protein…
Scientists have developed a new compression technique for protein databases that can be used in the Centrifuger method for taxonomic classification. This lossless compression algorithm improves upon the run-block compression method used for the Ferragina-Manzini (FM) index by scaling better with alphabet size. When tested on RefSeq prokaryotic and viral protein sequences, Centrifuger was found to reduce memory usage by more than a third compared to Kaiju, a method that builds on a plain FM-index.
Despite these improvements in memory efficiency, Centrifuger maintains comparable running time to Kaiju.
Another key advantage of Centrifuger is its ability to accurately locate matches of arbitrary length within the compressed FM-index. This feature allows the method to achieve higher accuracy in taxonomic classification than Kraken2, a k-mer-based approach. To demonstrate the practical applications of Centrifuger, researchers created an index of size 182 GB for classifying reads against the full nr database, which contains approximately 250 billion amino acid characters.
Using this index, Centrifuger was able to distinguish different SARS-CoV-2 infection states and viral transcriptome profiles across various human cell types in single-cell RNA-seq data.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.