Researchers Developed Protein Database Compression Method
A new algorithm enables taxonomic classification of massive protein datasets with significantly lower memory usage.
Updated on Sept. 24, 2026 in Biotech

Live Poll
Do you believe new data compression technologies will lead to significant improvements in public health research?
Scientists have created a new lossless compression algorithm for indexing protein databases that drastically reduces memory footprints. The method, called Centrifuger, facilitates faster and more accurate sequence classification compared to previous tools.
Why it matters
The development allows researchers to index massive biological datasets that were previously too memory-intensive to process efficiently. This advancement improves the capability of genomic tools to identify specific viral signatures and transcriptome profiles.
The new index occupies 182 GB of memory while classifying sequences against the nr database, which contains approximately 250 billion amino acid characters. The algorithm scales its efficiency based on the alphabet size of protein datasets.
The players
Centrifuger
This is the newly developed classification method designed for efficient taxonomic identification of protein sequences.
The details
The method employs a new run-block compression scheme to minimize the size of the FM-index. In testing, the tool demonstrated higher classification accuracy than the Kraken2 method and successfully identified SARS-CoV-2 infection states within human cell types.
Timeline
September 2026: The peer-reviewed paper describing the algorithm was published.
The Tech Race
The development of Centrifuger marks a significant shift from traditional indexing methods that struggle with the ballooning size of modern genomic databases. By optimizing memory usage, the tool challenges older standards like Kraken2 and sets a new benchmark for computational biology efficiency.
Researchers and developers can now utilize more compact database indexes, reducing the hardware requirements for large-scale genomic analysis. This improvement leads to faster processing times and more accessible computational pipelines for identifying viral transcriptomes.
The takeaway
This breakthrough provides a scalable solution to the computational bottleneck caused by the massive growth of protein databases. Implementing this algorithm could streamline diagnostics and disease surveillance efforts across the global biotech community.
Further reading
For more information on the evolving landscape of biological software tools, visit the Biotech section.
More information
Access the full research paper and algorithm details for technical specifications.
Source note: This article includes information reported by Biorxiv.
Live Poll
Do you believe new data compression technologies will lead to significant improvements in public health research?







