Researchers Closed Gap in 8-Bit LLM Training
A new method called Delta-Matching enables native 8-bit training that matches full-precision results.
Updated on Oct. 1, 2026 in Artificial Intelligence

Live Poll
Do you believe advancements in AI efficiency will significantly lower the cost of building large models?
Researchers from MIT and Carnegie Mellon University have identified the mathematical root cause for accuracy gaps in 8-bit large language model training. Their new method, Delta-Matching, restores training performance to levels seen in standard BF16 and FP32 baselines.
Why it matters
Native 8-bit training allows for significantly faster compute speeds on modern hardware, but previous attempts suffered from systemic gradient bias. This breakthrough overcomes the stale-delta effect to enable more efficient model development without sacrificing quality.
FP8 hardware on NVIDIA H100 GPUs offers 3,958 teraflops of compute performance, doubling the 1,979 teraflops available in BF16. The study tested these metrics on models with up to 5.29 billion parameters.
The players
MIT
This prestigious research university served as one of the institutional bases for the authors of the new training method.
Carnegie Mellon University
This institution collaborated on the study and provided key researchers for the development of Delta-Matching.
NVIDIA
This technology company produces the H100 GPUs that benefit from increased compute efficiency via FP8 training.
The details
The team discovered that standard FP8 quantization breaks the zero-row-sum invariant of the softmax Jacobian during backpropagation, causing scaling mismatches between forward and backward passes. Delta-Matching resolves this by adjusting scaling factors to restore these invariants without requiring architectural changes or reduced global batch sizes.
Timeline
In 2024, NAVER Cloud researchers documented irrecoverable FP8 training divergence.
On September 29, 2026, the Delta-Matching paper was posted to arXiv.
The Tech Race
This development represents a major leap in optimizing AI training for the NVIDIA H100 GPU architecture. By enabling native 8-bit workflows, it helps replace less efficient precision formats and positions future LLM training for massive reductions in compute time.
Developers and researchers can expect faster training times and lower hardware requirements for large language models. This shift may ultimately lead to more accessible and cost-effective AI model development cycles.
The takeaway
Solving mathematical quantization errors is a vital step toward making high-end AI development more sustainable. This research proves that precise adjustments to gradient scaling can yield massive improvements in training efficiency.
What happens next
The authors intend to release implementation code, trained checkpoints, and data recipes to allow for wider industry adoption and testing.
Further reading
For more on the evolution of model efficiency, visit our Artificial Intelligence section.
Source note: This article includes information reported by Tech Times.
Live Poll
Do you believe advancements in AI efficiency will significantly lower the cost of building large models?










