Prime Intellect Launched Prime Inference AI Platform
The company introduced a new serving platform that supports open-source models on NVIDIA Blackwell hardware.
Updated on Oct. 3, 2026 in Artificial Intelligence

Live Poll
Do you believe serverless platforms make building and scaling AI applications significantly easier for your projects?
Prime Intellect has launched Prime Inference, a new serving platform for frontier open-source models. The system utilizes company-owned NVIDIA Blackwell hardware across multiple datacenters to provide serverless endpoints.
Why it matters
The platform offers a specialized architecture for high-performance AI deployment that is compatible with OpenAI SDKs. By integrating advanced hardware and software, it provides optimized infrastructure for developers using large-scale language models.
The architecture achieves a 100% uptime rate and a 40% reduction in p90 inter-token latency by using prefill and decode disaggregation. The stack incorporates NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer to handle 1.63M cached tokens per decoder.
The players
Prime Intellect
Prime Intellect is a technology company that develops infrastructure and platforms for deploying large-scale artificial intelligence models.
NVIDIA
NVIDIA is a global technology corporation that designs graphics processing units and hardware platforms used for artificial intelligence and high-performance computing.
The details
The platform currently supports GLM-5.3 models on GB200 NVL72 hardware while maintaining compatibility with OpenAI-compatible SDKs. It employs cache-aware routing that uses host DRAM as a secondary KV tier to manage token processing for large prompts.
Timeline
The official platform launch announcement occurred on October 3, 2026.
Verification of comparative competitor pricing data took place on October 2, 2026.
The Tech Race
This release follows the ongoing industry transition toward hardware-accelerated serving stacks specifically built for the NVIDIA Blackwell GPU architecture. It places Prime Intellect in direct competition with cloud providers by offering dedicated, optimized inference paths for the newest generation of high-performance AI hardware.
Developers and companies using this platform can expect improved processing speeds for large prompts, with systems targeted to reach 100 tokens per second. The integration with OpenAI-compatible SDKs reduces the technical burden for teams migrating existing applications to this new inference stack.
The takeaway
The move demonstrates a shift toward highly optimized, hardware-specific inference stacks that allow for faster, more reliable model deployment. Developers should evaluate their current token throughput needs against the platform's cache-aware routing capabilities to determine if the infrastructure fits their workflow.
What happens next
Prime Intellect intends to expand the platform's capabilities by adding support for Vera Rubin hardware, as well as future features including batch inference and one-click dedicated deployments.
Further reading
Learn more about the latest industry developments at Artificial Intelligence.
Source note: This article includes information reported by MarkTechPost.
Live Poll
Do you believe serverless platforms make building and scaling AI applications significantly easier for your projects?










