Research — Tokenization

PU-Tok — An Indic-First Multilingual Tokenizer

A 200K-vocab tokenizer for 44 languages, built so non-Latin scripts don't pay the fertility penalty they take in standard BPE.

Production

The problem

Standard BPE tokenizers are trained on Latin-heavy corpora and it shows: Hindi, Tamil, or Bengali text routinely needs 2-4x more tokens per character than English for the same content. That's slower inference and worse context utilization for exactly the languages we care most about.

Approach

44
languages supported
200K
vocabulary size
38
test regression suite

Honest development notes

Not every idea worked. Morphology suffix protection caused a Hindi fertility regression before being fixed back to neutral. Iterative reallocation turned out to be a no-op on small corpora because the vocabulary was already saturated. Both are documented, not smoothed over.

Where PU-Tok still wins

Head-to-head against our own Hyper-Token on FLORES-200, PU-Tok comes out ahead on Cyrillic, Greek, Armenian, Thai, Lao, Khmer, and Amharic — scripts with dedicated handling in its pipeline.

© 2026 Cybergeon Technologies. All numbers on this site are re-verified against real code before publishing.