Clockwork.io Raised $31 Million for AI Fault Tolerance
The startup secured new capital to expand its software tools designed to minimize AI workload interruptions.
Updated on Oct. 5, 2026 in Artificial Intelligence

Live Poll
Do you believe investments in AI infrastructure reliability are necessary for modern business performance?
Clockwork.io has raised $31 million in a funding round co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures. This investment brings the company total funding to $73 million as it scales its AI fault-tolerance technology.
Why it matters
Distributed AI workloads are frequently plagued by hardware and network failures that cause significant GPU downtime. Clockwork.io offers software solutions to mitigate these disruptions and improve training efficiency.
The company reported that TorchPass reduced training goodput loss from 14% to 3% for neocloud customers. The software allows for recovery from checkpoint reloads that can take up to 90 minutes.
The players
Clockwork.io
This Palo Alto-based company specializes in fault-tolerance software for large-scale distributed AI infrastructure.
This professional networking platform utilizes LinkPass to protect its infrastructure fleet from connection failures.
Together AI
This organization offers GPU cluster services and provides TorchPass as a managed solution for its users.
WhiteFiber
This tech company is currently expanding the integration of Clockwork.io software across its global infrastructure footprint.
Premji Invest
This investment firm co-led the recent funding round for Clockwork.io alongside Wing Venture Capital and Seligman Ventures.
The details
Clockwork.io introduced multi-node platform snapshots and fast asynchronous application checkpoints for its TorchPass solution. The company software migrates training workloads from failing GPUs to healthy ones, while LinkPass reroutes network traffic around failed links to prevent job interruption.
Timeline
NEA first backed Clockwork.io in 2021.
Clockwork.io announced its new funding on October 5, 2026.
The Tech Race
As AI models require increasingly large GPU clusters like the 16,384 units used in the Meta Llama 3 training, fault tolerance has become a critical bottleneck for hardware utilization. This technology represents a shift from reactive hardware replacement to proactive software-defined resilience.
While these tools are designed for enterprise infrastructure, the improvements in training efficiency help lower the compute costs associated with developing large language models. Users may see faster AI service deployments and improved reliability from major cloud platforms integrated with this software.
The takeaway
Reliable infrastructure is essential for the future of large-scale artificial intelligence development. As model sizes grow, software solutions that enable rapid recovery from hardware failures will become standard in modern data centers.
Further reading
Learn more about the latest innovations in Artificial Intelligence.
More information
For more information, visit the Clockwork.io official company website.
Live Poll
Do you believe investments in AI infrastructure reliability are necessary for modern business performance?










