All posts
// / Blog

A client's chatbot was giving garbage responses to anything written in Hindi.

The model was fine — the tokenizer was the problem.

Here's what most engineers miss: the tokenizer is the foundation of every language model, and it's also the most common source of subtle bugs.

English text gets tokenized efficiently — common words become single tokens. But Hindi, Arabic, Chinese, and many other languages get fragmented into character-level pieces, burning through context windows and degrading quality.

The same model can be brilliant in English and mediocre in Hindi, purely because of tokenization.

Understanding tokenization means understanding: why "unhappiness" becomes ["un", "happi", "ness"] and not ["unhappy", "ness"]. Why your model handles some languages 3x less efficiently than others. Why special characters can break your pipeline. Why token limits aren't word limits.

Before I deploy any LLM application, I check three things: how does the tokenizer handle my target language? What's the actual token-to-word ratio for my data? Are there any token boundary issues in my use case?

It's not glamorous work. But it's saved me from shipping broken products more times than I can count.

#NLP#Tokenization#LLM#MultilingualAI#MachineLearning#AIEngineering