Researchers Closed Gap in 8-Bit LLM Training

A new method called Delta-Matching enables native 8-bit training that matches full-precision results.

Updated on Oct. 1, 2026 in Artificial Intelligence

Researchers Closed Gap in 8-Bit LLM Training

Live Poll

Do you believe advancements in AI efficiency will significantly lower the cost of building large models?

Researchers from MIT and Carnegie Mellon University have identified the mathematical root cause for accuracy gaps in 8-bit large language model training. Their new method, Delta-Matching, restores training performance to levels seen in standard BF16 and FP32 baselines.

Why it matters

Native 8-bit training allows for significantly faster compute speeds on modern hardware, but previous attempts suffered from systemic gradient bias. This breakthrough overcomes the stale-delta effect to enable more efficient model development without sacrificing quality.

FP8 hardware on NVIDIA H100 GPUs offers 3,958 teraflops of compute performance, doubling the 1,979 teraflops available in BF16. The study tested these metrics on models with up to 5.29 billion parameters.

The players

MIT

This prestigious research university served as one of the institutional bases for the authors of the new training method.

Carnegie Mellon University

This institution collaborated on the study and provided key researchers for the development of Delta-Matching.

NVIDIA

This technology company produces the H100 GPUs that benefit from increased compute efficiency via FP8 training.

The details

The team discovered that standard FP8 quantization breaks the zero-row-sum invariant of the softmax Jacobian during backpropagation, causing scaling mismatches between forward and backward passes. Delta-Matching resolves this by adjusting scaling factors to restore these invariants without requiring architectural changes or reduced global batch sizes.

Timeline

  1. In 2024, NAVER Cloud researchers documented irrecoverable FP8 training divergence.

  2. On September 29, 2026, the Delta-Matching paper was posted to arXiv.

The Tech Race

This development represents a major leap in optimizing AI training for the NVIDIA H100 GPU architecture. By enabling native 8-bit workflows, it helps replace less efficient precision formats and positions future LLM training for massive reductions in compute time.

Developers and researchers can expect faster training times and lower hardware requirements for large language models. This shift may ultimately lead to more accessible and cost-effective AI model development cycles.

The takeaway

Solving mathematical quantization errors is a vital step toward making high-end AI development more sustainable. This research proves that precise adjustments to gradient scaling can yield massive improvements in training efficiency.

What happens next

The authors intend to release implementation code, trained checkpoints, and data recipes to allow for wider industry adoption and testing.

Further reading

For more on the evolution of model efficiency, visit our Artificial Intelligence section.

Source note: This article includes information reported by Tech Times.

Live Poll

Do you believe advancements in AI efficiency will significantly lower the cost of building large models?