Researchers Solved LLM Character Counting Errors

A new byteification method allows models to accurately read individual characters within words.

Updated on Oct. 7, 2026 in Language Learning

Isometric editorial illustration of a segmented metallic data cube, representing precision and discrete byte structures in AI.
Researchers have developed a new byteification technique that allows large language models to process granular character data, effectively resolving common tokenization inaccuracies. AI Illustration. Upload story photo >

Live Poll

Do you trust artificial intelligence to perform simple tasks accurately in your daily life?

Many large language models struggle to identify letters within words, such as incorrectly counting only two r's in the word strawberry instead of three. Researchers have now developed a technique called byteification to enable these models to access individual characters.

Why it matters

Traditional token-based models encode sequences rather than letters, which limits their ability to process granular character data. This new approach retrofits existing models to operate at the byte level to resolve these common inaccuracies.

Most current LLMs encode words as tokens representing sequences of letters, whereas this new method retrofits them to encode individual characters as binary sequences called bytes.

The players

Minixhofer et al.

These researchers are responsible for developing the byteification method to improve LLM character recognition.

Nature

This prestigious scientific journal served as the publication platform for the new research findings.

The details

Developed by Minixhofer et al., the byteification approach allows models to bypass token limitations by operating at the byte level. This enables the technology to accurately identify and count specific characters that were previously obscured by standard tokenization.

Timeline

  1. October 7, 2026: The research was published.

The Big Picture

This research addresses a fundamental structural limitation inherent in current token-based LLM architectures. By enabling character-level access, the study marks a departure from the sequence-heavy processing that has dominated language model development.

Users will likely experience fewer logic errors in AI outputs regarding spelling, word structure, or character-specific queries. This shift reduces the need for manual fact-checking of AI-generated text for tasks requiring high precision in character counting.

The takeaway

Advancements in how AI processes base characters could significantly improve reliability for specialized tasks like coding or word games. Understanding these technical limitations helps users better anticipate when an AI model might struggle with granular linguistic detail.

Further reading

Learn more about the evolution of linguistic technology in our Language Learning section.

Source note: This article includes information reported by Nature.

Live Poll

Do you trust artificial intelligence to perform simple tasks accurately in your daily life?