Scality Released AI Inference Factory Software Stack

The new open-code platform aims to help enterprises manage on-premises AI inference with better control and scalability.

Updated on Oct. 8, 2026 in Data Centers

Bold flat-color editorial illustration of vertical server rack structures, representing enterprise data infrastructure.
Scality launched its AI Inference Factory, an open-code software stack designed to help enterprises manage AI model inference and data infrastructure locally. AI Illustration. Upload story photo >

Live Poll

Do you trust on-premises data infrastructure more than cloud-based AI services for your sensitive information?

Scality has launched its AI Inference Factory, an open-code software stack designed to support on-premises AI inference for enterprise users. The solution integrates a disaggregated inference-serving layer with Scality AI Data Infrastructure to provide better control over data sovereignty.

Why it matters

Organizations are increasingly seeking alternatives to unpredictable cloud-based AI costs while requiring more control over model lifecycles and data. This stack addresses those needs by providing a supported, scalable infrastructure for running models locally.

The system supports models like Mistral, Gemma, and DeepSeek, achieving a 1.9-second load time for Gemma-3 27B. Its architecture separates prefill from decode processes, supporting over 1,000 concurrent sessions on standard servers.

The players

Scality

Scality is a company that has spent 15 years building data infrastructure and storage solutions.

The details

The platform utilizes Scality AI Data Infrastructure as a shared cache layer, extending GPU memory via NVMe and RDMA to optimize performance. It enables 7.4 times more requests by separating tasks, while allowing enterprises to maintain the entire stack as either a license or managed service.

Timeline

  1. Scality announced and demonstrated the AI Inference Factory on October 8, 2026.

The Tech Race

This move signals a shift in the AI infrastructure wars, pitting on-premises control against the dominant cloud-based service model. By leveraging standard server hardware, Scality is attempting to replace proprietary cloud lock-in with a flexible, high-performance local stack.

Enterprises can now choose between software licenses or managed services to reduce dependency on public cloud AI providers. This allows IT teams to implement local AI inference models without the unpredictable costs often associated with cloud-based compute usage.

The takeaway

Bringing AI infrastructure on-premises can significantly mitigate the long-term volatility of cloud service pricing. Companies prioritizing data sovereignty should evaluate whether a disaggregated inference layer matches their current hardware footprint.

Further reading

For additional context on storage and infrastructure, explore the Data Centers section.

More information

Read more details on the Scality AI Inference Factory product page.

Live Poll

Do you trust on-premises data infrastructure more than cloud-based AI services for your sensitive information?