Vedavinya AI is
coming soon.
India's own proprietary foundation model — trained from scratch on 22 Indic languages and a curated corpus of thought, culture, science and code. We're putting the finishing touches on the private preview. Sovereign by design. Built in Bhārat, for the world.
Fourteen tokenizers,
forty languages.
A tokenizer decides how many pieces a sentence is cut into, and every piece is paid for — in cost, in latency, and in how much text fits in a model's memory. Indian languages are the ones the major tokenizers cut worst. On the seventy-ninth Independence Day, this is where ours stands against them, measured the same way for everyone, with what we lose shown alongside what we win.
Across 40 languages of real web text, VedaVinya-T needs 2.48 tokens per word. The next best tokenizer, Gemma-3 (Google), needs 2.63 — +5.9% more.
250 real documents per language from FineWeb-2, AI4Bharat Sangraha, FineWeb-Edu; 500 paired bootstrap replicates; all 14 tokenizers measured locally under one protocol, with no vendor-published figures mixed in.
Average fertility
Tokens per word — lower is better. 40 languages, real documents. This is the scope the significance tests were run on.
Percentages are measured against our own token count: +32% means that tokenizer needs a third more tokens than we do for the same text.
| Tokenizer | Built by | Origin | Tokens/word | vs VedaVinya-T | Round-trip |
|---|---|---|---|---|---|
| VedaVinya-T | VedaVinya | India | 2.48 | — | lossless |
| Gemma-3 | USA | 2.63 | +5.9% | lossless | |
| Gemma-2 | USA | 3.10 | +25.0% | lossless | |
| GPT-4o / OSS | OpenAI | USA | 3.28 | +32.3% | lossless |
| Sarvam-M | Sarvam AI | India | 3.85 | +55.2% | lossless |
| Nemotron-Nano | NVIDIA | USA | 3.85 | +55.2% | lossless |
| Sarvam-1 | Sarvam AI | India | 4.67 | +87.9% | lossless |
| Qwen2.5 | Alibaba | China | 5.52 | +122.1% | 97.44% — lossy |
| Qwen3 | Alibaba | China | 5.52 | +122.1% | 97.44% — lossy |
| Llama-3.1 | Meta | USA | 5.63 | +126.6% | 78.00% — lossy |
| GLM-4.5/4.6 | Zhipu AI | China | 6.22 | +150.5% | lossless |
| GPT-4 | OpenAI | USA | 6.44 | +159.3% | lossless |
| OLMo-2 | AI2 | USA | 6.44 | +159.3% | lossless |
| Grok-1 | xAI | USA | 10.12 | +307.4% | lossless |
Who built for which scripts
The same tokenizers, grouped by where they were built. Fourteen tokenizers from three countries.
India
3 tokenizers- Best
- 2.48 VedaVinya-T
- Group mean
- 3.67
- VedaVinya-Tours2.48
- Sarvam-M3.85
- Sarvam-14.67
USA
8 tokenizers- Best
- 2.63 Gemma-3
- Group mean
- 5.19
- Gemma-32.63
- Gemma-23.10
- GPT-4o / OSS3.28
- Nemotron-Nano3.85
- Llama-3.15.63
- GPT-46.44
- OLMo-26.44
- Grok-110.12
China
3 tokenizers- Best
- 5.52 Qwen2.5
- Group mean
- 5.75
- Qwen2.55.52
- Qwen35.52
- GLM-4.5/4.66.22
Read this as provenance, not nationality: the tokenizers that handle Indian scripts well are, with one close exception, the ones whose builders needed them to work.
Byte-exact round-trip
decode(encode(x)) == x, across all 42 languages. This gates everything above it.
Lossless — 100.00%
- VedaVinya-Tours
- Gemma-3
- Gemma-2
- GPT-4o / OSS
- Sarvam-M
- Nemotron-Nano
- Sarvam-1
- GLM-4.5/4.6
- GPT-4
- OLMo-2
- Grok-1
Not lossless
- Qwen2.597.44%worst: Manipuri (Sangraha) 32.8%
- Qwen397.44%worst: Manipuri (Sangraha) 32.8%
- Llama-3.178.00%worst: Santali (Sangraha) 35.2%
A tokenizer that cannot reconstruct the text can post a low token count by throwing input away, so this is checked before any fertility number is read.
Is the lead real?
Paired bootstrap in each of 40 languages, 500 replicates. A win means we use significantly fewer tokens; a tie means the interval straddles zero.
- Gemma-3274927·4·9+5.3%
- Gemma-2294729·4·7+24.1%
- GPT-4o / OSS33633·1·6+12.4%
- Sarvam-M3434·3·3+27.9%
- Nemotron-Nano3434·3·3+27.9%
- Sarvam-13939·1·0+49.6%
- Qwen2.53939·1·0+90.1%
- Qwen33939·1·0+90.1%
- Llama-3.13939·0·1+47.1%
- GLM-4.5/4.63838·0·2+129.8%
- GPT-43939·0·1+135.8%
- OLMo-23939·0·1+135.8%
- Grok-13939·1·0+230.1%
A tie is not a win. Gemma-3 ties 4 languages and beats us outright in 9 — that is a contest, and the page would be worth less if it hid it.
| Competitor | Win | Tie | Loss | Median gap |
|---|---|---|---|---|
| Gemma-3 | 27 | 4 | 9 | +5.3% |
| Gemma-2 | 29 | 4 | 7 | +24.1% |
| GPT-4o / OSS | 33 | 1 | 6 | +12.4% |
| Sarvam-M | 34 | 3 | 3 | +27.9% |
| Nemotron-Nano | 34 | 3 | 3 | +27.9% |
| Sarvam-1 | 39 | 1 | 0 | +49.6% |
| Qwen2.5 | 39 | 1 | 0 | +90.1% |
| Qwen3 | 39 | 1 | 0 | +90.1% |
| Llama-3.1 | 39 | 0 | 1 | +47.1% |
| GLM-4.5/4.6 | 38 | 0 | 2 | +129.8% |
| GPT-4 | 39 | 0 | 1 | +135.8% |
| OLMo-2 | 39 | 0 | 1 | +135.8% |
| Grok-1 | 39 | 1 | 0 | +230.1% |
Script coverage
Tokens per character on the bare alphabet — no corpus involved, so there is no sampling choice to argue about. Lower is better.
1.00 or above means the script is absent from the vocabulary and is being rebuilt byte by byte. 3.00 means no coverage at all. Meetei Mayek is our gap — and Gemma-2 is the only tokenizer here that really has it.
| Script | VedaVinya-T | Gemma-3 | Gemma-2 | GPT-4o / OSS | Sarvam-1 | Llama-3.1 | Grok-1 |
|---|---|---|---|---|---|---|---|
| Devanagari | 1.02 | 1.14 | 1.16 | 1.16 | 1.36 | 1.27 | 2.95 |
| Bengali | 1.14 | 1.21 | 1.28 | 1.23 | 1.51 | 1.98 | 3.02 |
| Tamil | 1.26 | 1.81 | 1.47 | 1.56 | 1.91 | 2.00 | 3.02 |
| Ol Chiki | 1.07 | 1.20 | 1.00 | 3.00 | 3.03 | 3.00 | 3.03 |
| Meetei Mayek | 2.69 | 2.50 | 1.31 | 3.00 | 3.03 | 3.00 | 3.03 |
| Gurmukhi | 1.21 | 1.42 | 1.37 | 1.30 | 1.40 | 2.00 | 3.02 |
| Odia | 1.19 | 1.42 | 1.51 | 1.53 | 1.58 | 2.98 | 3.02 |
| Malayalam | 1.05 | 1.19 | 1.23 | 1.26 | 1.56 | 2.00 | 3.02 |
Against the strongest current rival, script by script
For each script, the best current-generation tokenizer other than ours. Superseded models (GPT-4, Gemma-2, Qwen2.5) are excluded here.
- Latin13 languages1.60 vs 1.63 GPT-4o/OSS-2.0%
- Devanagari9 languages2.31 vs 2.39 Gemma-3-3.2%
- Perso-Arabic4 languages1.66 vs 1.76 GPT-4o/OSS-5.6%
- Bengali-Assamese4 languages2.69 vs 2.68 Gemma-3+0.5%
- Ol Chiki2 languages5.33 vs 5.40 Gemma-3-1.3%
- Gujarati1 language1.77 vs 2.02 Sarvam-1-12.0%
- Kannada1 language2.53 vs 2.51 Sarvam-1+0.8%
- Hangul1 language2.66 vs 2.60 Nemotron-Nano+2.2%
- Malayalam1 language3.05 vs 3.34 Sarvam-1-8.6%
- Meetei Mayek1 language14.77 vs 11.91 Gemma-3+24.1%
- Odia1 language1.97 vs 2.47 Sarvam-1-20.4%
- Gurmukhi1 language1.63 vs 1.69 Sarvam-1-3.8%
- Cyrillic1 language2.13 vs 2.00 GPT-4o/OSS+6.9%
- Tamil1 language2.17 vs 2.44 Gemma-3-11.0%
- Telugu1 language2.29 vs 2.47 Sarvam-1-7.2%
Negative means we use fewer tokens. Three of fifteen scripts go against us — and one of them, Meetei Mayek, goes against us badly.
Every language
All 42 rows, searchable. Pick the tokenizers to compare against.
| Language | VedaVinya-T | Gemma-3 | GPT-4o / OSS | Sarvam-1 | Llama-3.1 |
|---|---|---|---|---|---|
| Arabic | 1.90 | 2.06 | 1.89 | 10.48 | 2.28✳ |
| Assamese | 2.20 | 2.64 | 2.62 | 4.13 | 8.03✳ |
| Bengali | 1.71 | 1.68 | 2.31 | 2.03 | 7.76✳ |
| Bodo | 2.81 | 2.98 | 3.34 | 3.07 | 3.80✳ |
| Dogri | 2.08 | 1.95 | 2.15 | 2.13 | 2.82✳ |
| English | 1.39 | 1.39 | 1.34 | 1.62 | 1.35✳ |
| French | 1.53 | 1.55 | 1.47 | 2.65 | 1.72✳ |
| German | 1.77 | 1.75 | 1.67 | 3.20 | 2.00✳ |
| Gujarati | 1.77 | 2.35 | 2.25 | 2.02 | 9.36✳ |
| Hindi | 1.28 | 1.37 | 1.61 | 1.45 | 2.64✳ |
| Hindi (romanised) | 1.76 | 2.11 | 2.07 | 2.92 | 2.31✳ |
| Hinglish | 1.38 | 1.52 | 1.61 | 2.31 | 1.78✳ |
| Indonesian | 1.62 | 1.63 | 1.69 | 3.09 | 2.08✳ |
| Italian | 1.58 | 1.61 | 1.65 | 2.62 | 1.86✳ |
| Kannada | 2.53 | 3.14 | 3.08 | 2.51 | 13.52✳ |
| Kashmiri | 1.59 | 1.64 | 1.73 | 6.61 | 2.20✳ |
| Konkani | 2.45 | 2.63 | 2.77 | 2.99 | 3.55✳ |
| Korean | 2.66 | 2.62 | 2.70 | 10.02 | 2.76✳ |
| Maithili | 1.98 | 2.10 | 2.19 | 2.27 | 3.02✳ |
| Malayalam | 3.05 | 3.47 | 3.50 | 3.34 | 15.84✳ |
| Manipuri | 14.77 | 11.91 | 15.82 | 16.72 | 15.82✳ |
| Manipuri (Bengali) | 3.77 | 3.51 | 4.21 | 4.00 | 9.35✳ |
| Manipuri (Sangraha) | 3.08 | 2.88 | 3.40 | 3.54 | 7.02✳ |
| Marathi | 1.74 | 1.95 | 2.49 | 1.90 | 3.82✳ |
| Mixed† | 1.91 | 1.64 | 1.45 | 3.82 | 1.73 |
| Nepali | 1.79 | 2.07 | 2.15 | 2.41 | 3.52✳ |
| Odia | 1.97 | 4.35 | 6.09 | 2.47 | 14.68✳ |
| Panini† | 3.40 | 3.20 | 3.40 | 3.20 | 4.20 |
| Portuguese | 1.53 | 1.52 | 1.44 | 2.73 | 1.72✳ |
| Punjabi | 1.63 | 2.69 | 2.60 | 1.69 | 7.69✳ |
| Russian | 2.13 | 2.00 | 2.00 | 12.67 | 2.28✳ |
| Sanskrit | 3.27 | 3.23 | 3.70 | 3.59 | 4.54✳ |
| Santali | 5.36 | 5.43 | 13.72 | 13.77 | 12.88✳ |
| Santali (Sangraha) | 5.31 | 5.37 | 13.56 | 13.61 | 12.69✳ |
| Sindhi | 1.68 | 2.16 | 1.84 | 8.08 | 2.96✳ |
| Spanish | 1.47 | 1.42 | 1.39 | 2.57 | 1.62✳ |
| Swahili | 1.63 | 2.13 | 1.87 | 3.19 | 2.44✳ |
| Tamil | 2.17 | 2.44 | 3.16 | 2.48 | 12.10✳ |
| Telugu | 2.29 | 2.96 | 3.17 | 2.47 | 13.00✳ |
| Turkish | 2.00 | 2.17 | 2.21 | 4.73 | 2.20✳ |
| Urdu | 1.48 | 1.50 | 1.58 | 8.08 | 2.87✳ |
| Vietnamese | 1.21 | 1.27 | 1.35 | 4.49 | 1.27✳ |
† marks a constructed single-sentence probe, excluded from the 40-language statistics. ✳ marks a language where that tokenizer is not byte-exact, so its token count is not a like-for-like score.
How this was measured
- 250 real documents per language, drawn from FineWeb-2, AI4Bharat Sangraha and FineWeb-Edu — two independent organisations — plus SentiMix for code-mixed text.
- 500 paired bootstrap replicates per language for significance, seed 20260815.
- All 14 tokenizers run locally under one protocol, on the same documents, in the same process. No vendor-published numbers are mixed in.
- Fertility is tokens per word, so lower is better. Round-trip is checked first: a score from a tokenizer that cannot reproduce the text is not a score.
What we are not claiming
Gemma-3 is close
+5.9% overall, and we lose 9 of 40 languages to it outright with 4 more too close to call. It is a real contest, not a clean sweep — and saying so is what makes the 39–1–0 against Sarvam-1 worth believing.
Grok-1 is xAI's 2024 tokenizer
It is the only tokenizer xAI has released. It should not be read as the tokenizer behind any current Grok model, and we do not label it "Grok".
Our eval and training share a source for Indic
AI4Bharat Sangraha appears in both. FineWeb-2 is the primary source everywhere it exists in order to bound that overlap, but it is a real limitation and we would rather state it than have it found.
Two scopes, and only one carries statistics
The headline 2.48 is 40 real-document languages — the scope the bootstrap ran on. A 42-language figure of 2.49 also exists which adds two single-sentence constructed probes; it is comparable to the world scorecard but not to the other scopes, and no significance test stands behind the probe rows.
Tokens per word needs words
Chinese, Japanese and Thai were collected and tokenized but are excluded from the fertility tables: they are not space-delimited, so "tokens per word" is not a meaningful measure for them.
This page publishes results and method. The tokenizer's vocabulary, merge tables and model weights are not published or available for download, and there is no tokenize-your-own-text box — token counts for chosen strings would leak the vocabulary one query at a time.