Researchers Developed Protein Database Compression Method

A new algorithm enables taxonomic classification of massive protein datasets with significantly lower memory usage.

Updated on Sept. 24, 2026 in Biotech

Researchers Developed Protein Database Compression Method

Live Poll

Do you believe new data compression technologies will lead to significant improvements in public health research?

Scientists have created a new lossless compression algorithm for indexing protein databases that drastically reduces memory footprints. The method, called Centrifuger, facilitates faster and more accurate sequence classification compared to previous tools.

Why it matters

The development allows researchers to index massive biological datasets that were previously too memory-intensive to process efficiently. This advancement improves the capability of genomic tools to identify specific viral signatures and transcriptome profiles.

The new index occupies 182 GB of memory while classifying sequences against the nr database, which contains approximately 250 billion amino acid characters. The algorithm scales its efficiency based on the alphabet size of protein datasets.

The players

Centrifuger

This is the newly developed classification method designed for efficient taxonomic identification of protein sequences.

The details

The method employs a new run-block compression scheme to minimize the size of the FM-index. In testing, the tool demonstrated higher classification accuracy than the Kraken2 method and successfully identified SARS-CoV-2 infection states within human cell types.

Timeline

  1. September 2026: The peer-reviewed paper describing the algorithm was published.

The Tech Race

The development of Centrifuger marks a significant shift from traditional indexing methods that struggle with the ballooning size of modern genomic databases. By optimizing memory usage, the tool challenges older standards like Kraken2 and sets a new benchmark for computational biology efficiency.

Researchers and developers can now utilize more compact database indexes, reducing the hardware requirements for large-scale genomic analysis. This improvement leads to faster processing times and more accessible computational pipelines for identifying viral transcriptomes.

The takeaway

This breakthrough provides a scalable solution to the computational bottleneck caused by the massive growth of protein databases. Implementing this algorithm could streamline diagnostics and disease surveillance efforts across the global biotech community.

Further reading

For more information on the evolving landscape of biological software tools, visit the Biotech section.

More information

Access the full research paper and algorithm details for technical specifications.

Source note: This article includes information reported by Biorxiv.

Live Poll

Do you believe new data compression technologies will lead to significant improvements in public health research?