A client's chatbot was giving garbage responses to anything written in Hindi.
The model was fine — the tokenizer was the problem.
Here's what most engineers miss: the tokenizer is the foundation of every language model, and it's also the most common source of subtle bugs.
English text gets tokenized efficiently — common words become single tokens. But Hindi, Arabic, Chinese, and many other languages get fragmented into character-level pieces, burning through context windows and degrading quality.
The same model can be brilliant in English and mediocre in Hindi, purely because of tokenization.
Understanding tokenization means understanding: why "unhappiness" becomes ["un", "happi", "ness"] and not ["unhappy", "ness"]. Why your model handles some languages 3x less efficiently than others. Why special characters can break your pipeline. Why token limits aren't word limits.
Before I deploy any LLM application, I check three things: how does the tokenizer handle my target language? What's the actual token-to-word ratio for my data? Are there any token boundary issues in my use case?
It's not glamorous work. But it's saved me from shipping broken products more times than I can count.