Providers Offered Hosted Access to Kimi K3 Model
Modal, Fireworks AI, and Baseten have launched hosting services for Moonshot AI's massive 2.8 trillion parameter model.
Updated on Sept. 21, 2026 in Artificial Intelligence

Live Poll
Would you trust a third-party provider to host your company's sensitive data for AI processing?
Modal, Fireworks AI, and Baseten have begun offering hosted inference services for Moonshot AI's Kimi K3 model. These providers allow enterprise users to access the massive AI model at a fraction of the cost of direct routes.
Why it matters
US export controls currently limit Chinese firms' access to high-end Nvidia and AMD hardware, while enterprise clients require US-hosted endpoints to maintain data sovereignty compliance.
Kimi K3 features 2.8 trillion parameters and an 896-expert architecture, activating 104 billion parameters per query. Hosted versions support a 1 million token context window and weights exceeding 1.4 terabytes.
The players
Moonshot AI
This Beijing-based artificial intelligence startup is the developer of the Kimi large language model series.
Modal
Modal is a serverless cloud computing provider that hosts machine learning models for developers and enterprises.
Fireworks AI
Fireworks AI is a platform designed for the high-speed deployment and inference of open-weights generative AI models.
Baseten
Baseten provides infrastructure tools that enable companies to deploy and scale AI models in production environments.
The details
Providers leverage Nvidia GB300 and AMD MI350X/MI355X accelerators, with Modal achieving speeds of 460 tokens per second using its DFlash speculative decoder. The services feature zero-data retention policies and offer OpenAI-compatible APIs for seamless integration.
Timeline
July 27, 2026: Moonshot AI released the Kimi K3 model.
The Tech Race
The emergence of these hosting services marks a shift in the AI arms race where cloud infrastructure providers compete to unlock access to globally distributed, high-parameter models. This approach circumvents hardware procurement challenges for developers while standardizing model access via common API interfaces.
Developers and enterprises can now utilize high-performance AI models through cost-effective APIs rather than managing expensive internal hardware stacks. These hosted endpoints offer businesses a clear path to integrate advanced AI capabilities while ensuring data sovereignty compliance.
The takeaway
The move to host massive foreign models on domestic cloud infrastructure illustrates how providers are bridging the gap between hardware restrictions and enterprise software demand. Companies looking to implement large-scale AI should prioritize providers that offer both high-speed inference and verifiable data privacy protocols.
Further reading
For more information on the current landscape of model deployment, visit our Artificial Intelligence section.
Live Poll
Would you trust a third-party provider to host your company's sensitive data for AI processing?










