Vedavinya AI · Private preview · 2026

Vedavinya AI is
coming soon.

India's own proprietary foundation model — trained from scratch on 22 Indic languages and a curated corpus of thought, culture, science and code. We're putting the finishing touches on the private preview. Sovereign by design. Built in Bhārat, for the world.

VedaVinya-T · Tokenizer benchmark · 15 August 2026

Fourteen tokenizers,
forty languages.

A tokenizer decides how many pieces a sentence is cut into, and every piece is paid for — in cost, in latency, and in how much text fits in a model's memory. Indian languages are the ones the major tokenizers cut worst. On the seventy-ninth Independence Day, this is where ours stands against them, measured the same way for everyone, with what we lose shown alongside what we win.

2.48
tokens per word

Across 40 languages of real web text, VedaVinya-T needs 2.48 tokens per word. The next best tokenizer, Gemma-3 (Google), needs 2.63 — +5.9% more.

250 real documents per language from FineWeb-2, AI4Bharat Sangraha, FineWeb-Edu; 500 paired bootstrap replicates; all 14 tokenizers measured locally under one protocol, with no vendor-published figures mixed in.

100.00% byte-exact round-trip40 languages · 14 tokenizersRun 15 August 2026
Language set
Built in
Colour
01

Average fertility

Tokens per word — lower is better. 40 languages, real documents. This is the scope the significance tests were run on.

Percentages are measured against our own token count: +32% means that tokenizer needs a third more tokens than we do for the same text.

02

Who built for which scripts

The same tokenizers, grouped by where they were built. Fourteen tokenizers from three countries.

India

3 tokenizers
Best
2.48 VedaVinya-T
Group mean
3.67
  • VedaVinya-Tours2.48
  • Sarvam-M3.85
  • Sarvam-14.67

USA

8 tokenizers
Best
2.63 Gemma-3
Group mean
5.19
  • Gemma-32.63
  • Gemma-23.10
  • GPT-4o / OSS3.28
  • Nemotron-Nano3.85
  • Llama-3.15.63
  • GPT-46.44
  • OLMo-26.44
  • Grok-110.12

China

3 tokenizers
Best
5.52 Qwen2.5
Group mean
5.75
  • Qwen2.55.52
  • Qwen35.52
  • GLM-4.5/4.66.22

Read this as provenance, not nationality: the tokenizers that handle Indian scripts well are, with one close exception, the ones whose builders needed them to work.

04

Byte-exact round-trip

decode(encode(x)) == x, across all 42 languages. This gates everything above it.

Lossless — 100.00%

  • VedaVinya-Tours
  • Gemma-3
  • Gemma-2
  • GPT-4o / OSS
  • Sarvam-M
  • Nemotron-Nano
  • Sarvam-1
  • GLM-4.5/4.6
  • GPT-4
  • OLMo-2
  • Grok-1

Not lossless

  • Qwen2.597.44%worst: Manipuri (Sangraha) 32.8%
  • Qwen397.44%worst: Manipuri (Sangraha) 32.8%
  • Llama-3.178.00%worst: Santali (Sangraha) 35.2%

A tokenizer that cannot reconstruct the text can post a low token count by throwing input away, so this is checked before any fertility number is read.

03

Is the lead real?

Paired bootstrap in each of 40 languages, 500 replicates. A win means we use significantly fewer tokens; a tie means the interval straddles zero.

we use fewer tie — too close to call they use fewer
  • Gemma-3274927·4·9+5.3%
  • Gemma-2294729·4·7+24.1%
  • GPT-4o / OSS33633·1·6+12.4%
  • Sarvam-M3434·3·3+27.9%
  • Nemotron-Nano3434·3·3+27.9%
  • Sarvam-13939·1·0+49.6%
  • Qwen2.53939·1·0+90.1%
  • Qwen33939·1·0+90.1%
  • Llama-3.13939·0·1+47.1%
  • GLM-4.5/4.63838·0·2+129.8%
  • GPT-43939·0·1+135.8%
  • OLMo-23939·0·1+135.8%
  • Grok-13939·1·0+230.1%

A tie is not a win. Gemma-3 ties 4 languages and beats us outright in 9 — that is a contest, and the page would be worth less if it hid it.

05

Script coverage

Tokens per character on the bare alphabet — no corpus involved, so there is no sampling choice to argue about. Lower is better.

VedaVinya-T
Gemma-3
Gemma-2
GPT-4o / OSS
Sarvam-1
Llama-3.1
Grok-1
Devanagari
best1.02
1.14
1.16
1.16
1.36
1.27
2.95
Bengali
best1.14
1.21
1.28
1.23
1.51
1.98
3.02
Tamil
best1.26
1.81
1.47
1.56
1.91
2.00
3.02
Ol ChikiSantali
1.07
1.20
best1.00
3.00
3.03
3.00
3.03
Meetei MayekManipuri
2.69
2.50
best1.31
3.00
3.03
3.00
3.03
Gurmukhi
best1.21
1.42
1.37
1.30
1.40
2.00
3.02
Odia
best1.19
1.42
1.51
1.53
1.58
2.98
3.02
Malayalam
best1.05
1.19
1.23
1.26
1.56
2.00
3.02
best lowest on that scriptin vocabularybyte fallback

1.00 or above means the script is absent from the vocabulary and is being rebuilt byte by byte. 3.00 means no coverage at all. Meetei Mayek is our gap — and Gemma-2 is the only tokenizer here that really has it.

07

Against the strongest current rival, script by script

For each script, the best current-generation tokenizer other than ours. Superseded models (GPT-4, Gemma-2, Qwen2.5) are excluded here.

  • Latin13 languages1.60 vs 1.63 GPT-4o/OSS-2.0%
  • Devanagari9 languages2.31 vs 2.39 Gemma-3-3.2%
  • Perso-Arabic4 languages1.66 vs 1.76 GPT-4o/OSS-5.6%
  • Bengali-Assamese4 languages2.69 vs 2.68 Gemma-3+0.5%
  • Ol Chiki2 languages5.33 vs 5.40 Gemma-3-1.3%
  • Gujarati1 language1.77 vs 2.02 Sarvam-1-12.0%
  • Kannada1 language2.53 vs 2.51 Sarvam-1+0.8%
  • Hangul1 language2.66 vs 2.60 Nemotron-Nano+2.2%
  • Malayalam1 language3.05 vs 3.34 Sarvam-1-8.6%
  • Meetei Mayek1 language14.77 vs 11.91 Gemma-3+24.1%
  • Odia1 language1.97 vs 2.47 Sarvam-1-20.4%
  • Gurmukhi1 language1.63 vs 1.69 Sarvam-1-3.8%
  • Cyrillic1 language2.13 vs 2.00 GPT-4o/OSS+6.9%
  • Tamil1 language2.17 vs 2.44 Gemma-3-11.0%
  • Telugu1 language2.29 vs 2.47 Sarvam-1-7.2%

Negative means we use fewer tokens. Three of fifteen scripts go against us — and one of them, Meetei Mayek, goes against us badly.

06

Every language

All 42 rows, searchable. Pick the tokenizers to compare against.

42 shown
Compare with
Tokens per word for each language
LanguageVedaVinya-TGemma-3GPT-4o / OSSSarvam-1Llama-3.1
Arabic1.902.061.8910.482.28✳
Assamese2.202.642.624.138.03✳
Bengali1.711.682.312.037.76✳
Bodo2.812.983.343.073.80✳
Dogri2.081.952.152.132.82✳
English1.391.391.341.621.35✳
French1.531.551.472.651.72✳
German1.771.751.673.202.00✳
Gujarati1.772.352.252.029.36✳
Hindi1.281.371.611.452.64✳
Hindi (romanised)1.762.112.072.922.31✳
Hinglish1.381.521.612.311.78✳
Indonesian1.621.631.693.092.08✳
Italian1.581.611.652.621.86✳
Kannada2.533.143.082.5113.52✳
Kashmiri1.591.641.736.612.20✳
Konkani2.452.632.772.993.55✳
Korean2.662.622.7010.022.76✳
Maithili1.982.102.192.273.02✳
Malayalam3.053.473.503.3415.84✳
Manipuri14.7711.9115.8216.7215.82✳
Manipuri (Bengali)3.773.514.214.009.35✳
Manipuri (Sangraha)3.082.883.403.547.02✳
Marathi1.741.952.491.903.82✳
Mixed†1.911.641.453.821.73
Nepali1.792.072.152.413.52✳
Odia1.974.356.092.4714.68✳
Panini†3.403.203.403.204.20
Portuguese1.531.521.442.731.72✳
Punjabi1.632.692.601.697.69✳
Russian2.132.002.0012.672.28✳
Sanskrit3.273.233.703.594.54✳
Santali5.365.4313.7213.7712.88✳
Santali (Sangraha)5.315.3713.5613.6112.69✳
Sindhi1.682.161.848.082.96✳
Spanish1.471.421.392.571.62✳
Swahili1.632.131.873.192.44✳
Tamil2.172.443.162.4812.10✳
Telugu2.292.963.172.4713.00✳
Turkish2.002.172.214.732.20✳
Urdu1.481.501.588.082.87✳
Vietnamese1.211.271.354.491.27✳

† marks a constructed single-sentence probe, excluded from the 40-language statistics. ✳ marks a language where that tokenizer is not byte-exact, so its token count is not a like-for-like score.

How this was measured

  • 250 real documents per language, drawn from FineWeb-2, AI4Bharat Sangraha and FineWeb-Edu — two independent organisations — plus SentiMix for code-mixed text.
  • 500 paired bootstrap replicates per language for significance, seed 20260815.
  • All 14 tokenizers run locally under one protocol, on the same documents, in the same process. No vendor-published numbers are mixed in.
  • Fertility is tokens per word, so lower is better. Round-trip is checked first: a score from a tokenizer that cannot reproduce the text is not a score.

What we are not claiming

Gemma-3 is close

+5.9% overall, and we lose 9 of 40 languages to it outright with 4 more too close to call. It is a real contest, not a clean sweep — and saying so is what makes the 39–1–0 against Sarvam-1 worth believing.

Grok-1 is xAI's 2024 tokenizer

It is the only tokenizer xAI has released. It should not be read as the tokenizer behind any current Grok model, and we do not label it "Grok".

Our eval and training share a source for Indic

AI4Bharat Sangraha appears in both. FineWeb-2 is the primary source everywhere it exists in order to bound that overlap, but it is a real limitation and we would rather state it than have it found.

Two scopes, and only one carries statistics

The headline 2.48 is 40 real-document languages — the scope the bootstrap ran on. A 42-language figure of 2.49 also exists which adds two single-sentence constructed probes; it is comparable to the world scorecard but not to the other scopes, and no significance test stands behind the probe rows.

Tokens per word needs words

Chinese, Japanese and Thai were collected and tokenized but are excluded from the fertility tables: they are not space-delimited, so "tokens per word" is not a meaningful measure for them.

This page publishes results and method. The tokenizer's vocabulary, merge tables and model weights are not published or available for download, and there is no tokenize-your-own-text box — token counts for chosen strings would leak the vocabulary one query at a time.

1B
Parameters · V1 target
22
Indic languages, day one
100%
Sovereign training data
Public
Model card & evals