Researchers Released New Microscopy Benchmark HSD-Bench
The dataset exposes performance inflation in machine learning models used to analyze blood cell images.
Updated on Oct. 8, 2026 in Life Sciences

Live Poll
Do you trust the accuracy claims of current artificial intelligence systems?
Scientists have introduced HSD-Bench, a new microscopy corpus containing 129,173 images across 14 blood cell classes. The benchmark highlights how improper data partitioning leads to significant performance inflation in machine learning models.
Why it matters
Machine learning benchmarks often treat related images as independent, which creates misleading performance metrics. This new tool forces researchers to account for source structure to ensure model accuracy reflects real-world capabilities.
The HSD-Bench corpus covers 14 distinct blood cell and hematopoietic cell states. Tests showed that random assignment of individual files pushed Vision Transformer accuracy from 0.851 to 0.990.
The players
HSD-Bench
This is a specialized microscopy benchmark designed to improve the evaluation of machine learning models in hematology.
The details
By reconstructing analysis groups using source filenames, researchers demonstrated that performance gains are often artificial artifacts of testing models on data that mirrors the training set. These inflation patterns were most pronounced within immature granulocytic and blast cell classes.
Timeline
October 8, 2026: The research paper and HSD-Bench dataset were released.
The Big Picture
This study updates the strictness of the ImageNet Large Scale Visual Recognition Challenge benchmarking standards to prevent data leakage in biomedical imaging. The work forces a paradigm shift in how biological vision models are validated against real-world diagnostic complexity.
This benchmark could lead to more reliable diagnostic tools by preventing AI models from appearing more accurate than they truly are. Standardizing these testing methods may accelerate the clinical adoption of automated hematology analysis.
The takeaway
Researchers must prioritize rigorous data partitioning over raw model output to ensure scientific validity. This study serves as a reminder that metrics in machine learning are only as reliable as the methods used to test them.
Further reading
Learn more about the latest innovations in Life Sciences.
More information
Read the complete research paper and HSD-Bench data on the bioRxiv platform.
Source note: This article includes information reported by Biorxiv.
Live Poll
Do you trust the accuracy claims of current artificial intelligence systems?







