pipette
ENEnglish

Lossless compression of protein databases for efficient and accurate metagenomic sequence classification with Centrifuger

L. Song

PreprintUso en el mundo real

En palabras de los autores

We present a lossless compression algorithm for indexing a protein database while supporting fast taxonomic classification in the method Centrifuger. The algorithm is a new scheme of the previously proposed run-block compression algorithm to reduce the size of the Ferragina-Manzini (FM) index, and it scales better with alphabet size than the original. On the RefSeq prokaryotic and viral protein sequences, Centrifuger reduces the memory footprint by over a third compared to the method Kaiju that builds on a plain FM-index, while having comparable running time. Furthermore, the compressed FM-index is lossless and can locate matches of arbitrary length, which helps Centrifuger achieve greater accuracy than Kraken2, a k-mer-based taxonomic classification method. We leverage the computational efficiency of Centrifuger to create an index of size 182 GB for classifying the reads against the full nr database that contains about 250 billion amino acid characters. Using this index, Centrifuger reveals different SARS-CoV-2 infection states and viral transcriptome profiles across human cell types from single-cell RNA-seq data.

Resultado principalEl resumen no menciona limitaciones.

Apareció: viernes, 25 de septiembre. bioRxiv. Preprint, todavía sin revisión por pares.

DOI: 10.64898/2026.09.17.752512