Clockwork.io Raised $31 Million for AI Fault Tolerance

The startup secured new capital to expand its software tools designed to minimize AI workload interruptions.

Updated on Oct. 5, 2026 in Artificial Intelligence

Bold vector editorial illustration of stacked server hardware and network cabling, representing resilient computing infrastructure.
Clockwork.io raised $31 million in a funding round led by Premji Invest to scale its AI fault-tolerance technology for enterprise workloads. AI Illustration. Upload story photo >

Live Poll

Do you believe investments in AI infrastructure reliability are necessary for modern business performance?

Clockwork.io has raised $31 million in a funding round co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures. This investment brings the company total funding to $73 million as it scales its AI fault-tolerance technology.

Why it matters

Distributed AI workloads are frequently plagued by hardware and network failures that cause significant GPU downtime. Clockwork.io offers software solutions to mitigate these disruptions and improve training efficiency.

The company reported that TorchPass reduced training goodput loss from 14% to 3% for neocloud customers. The software allows for recovery from checkpoint reloads that can take up to 90 minutes.

The players

Clockwork.io

This Palo Alto-based company specializes in fault-tolerance software for large-scale distributed AI infrastructure.

LinkedIn

This professional networking platform utilizes LinkPass to protect its infrastructure fleet from connection failures.

Together AI

This organization offers GPU cluster services and provides TorchPass as a managed solution for its users.

WhiteFiber

This tech company is currently expanding the integration of Clockwork.io software across its global infrastructure footprint.

Premji Invest

This investment firm co-led the recent funding round for Clockwork.io alongside Wing Venture Capital and Seligman Ventures.

The details

Clockwork.io introduced multi-node platform snapshots and fast asynchronous application checkpoints for its TorchPass solution. The company software migrates training workloads from failing GPUs to healthy ones, while LinkPass reroutes network traffic around failed links to prevent job interruption.

Timeline

  1. NEA first backed Clockwork.io in 2021.

  2. Clockwork.io announced its new funding on October 5, 2026.

The Tech Race

As AI models require increasingly large GPU clusters like the 16,384 units used in the Meta Llama 3 training, fault tolerance has become a critical bottleneck for hardware utilization. This technology represents a shift from reactive hardware replacement to proactive software-defined resilience.

While these tools are designed for enterprise infrastructure, the improvements in training efficiency help lower the compute costs associated with developing large language models. Users may see faster AI service deployments and improved reliability from major cloud platforms integrated with this software.

The takeaway

Reliable infrastructure is essential for the future of large-scale artificial intelligence development. As model sizes grow, software solutions that enable rapid recovery from hardware failures will become standard in modern data centers.

Further reading

Learn more about the latest innovations in Artificial Intelligence.

More information

For more information, visit the Clockwork.io official company website.

Live Poll

Do you believe investments in AI infrastructure reliability are necessary for modern business performance?