Tamil Pronunciation · Full-Alphabet Model Scorecard

Every Tamil letter, scored

All 30 core letters — 18 consonants and 12 vowels — recognised across speakers it has never heard, by a four-model speech ensemble we built.

61%30-way accuracy, unseen speakers (chance ≈ 3%)
94%mean vowel recognition @10% FAR
79%mean consonant recognition @10% FAR
21speakers · leave-one-speaker-out
01

How we got here — the model ladder

Recognition of ழ், Tamil's hardest sound, at each step. Every rung is a more Tamil-aware model.

Live engine
(formants + CNN)
8%
ழ் recog.
Multilingual
(128-lang)
42%
ழ்
18-way 52%
AI4Bharat
IndicWav2Vec
53%
ழ்
18-way 59%
Tamil-phoneme
(sound-trained)
75%
ழ்
18-way 55%
Ensemble
(all four)
67%
ழ் · balanced
30-way 61%
The arc

ழ் recognition climbs 8% → 42% → 53% → 75% as models get more Tamil-aware — but no single model wins everything. The system we built is an ensemble that fuses AI4Bharat's IndicWav2Vec with three other pretrained encoders; it trades a little peak accuracy for balance across all 30 letters.

02

Consonants — மெய் (18)

Cross-speaker recognition (TPR) at a fixed 10% false-accept budget. The small grey number is 30-way cross-letter accuracy — how often the letter beats all 29 others outright.

reliable (≥80%) workable (60–79%) hard (<60%)
TPR @10% FARTPR30-way
Consonant mean: 79% TPR@10% FAR  ·  the residual is the retroflex/alveolar sonorant cluster
03

Vowels — உயிர் (12)

Tamil vowels are acoustically distinct — the ensemble separates them almost perfectly.

TPR @10% FARTPR30-way
Vowel mean: 94% TPR@10% FAR  ·  several at 100% (ஆ ஐ ஈ உ) — effectively solved
04

What's solved, what's left

Solved

Vowels (94%) and distinct-place consonantsச் க் ஞ் த் ப் ம் ட் all 83–98%. The pretrained ensemble handles the bulk of the alphabet cross-speaker.

The residual — one nameable cluster

The retroflex & alveolar sonorants — ண் ன் ர் ற் ல் ள் ழ் — 53–78%. The tightest single confusion is ள்ல், the two laterals. This is exactly where more speaker-diverse data or cluster-specific fine-tuning pays off.

05

The errors, decoded

Where the residual confusion actually lands — and why it's phonetically principled, not random.

ழ் / ள் / ல் — 3-way confusion (counts · rows = true, columns = predicted)
true ↓ / pred →ழ்ள்ல்recall
ழ்2310364%
ள்11111431%
ல்7171233%

ழ் is the strong one (64%). The real mush is ள்ல் — the two laterals bleeding into each other almost symmetrically (14 and 17). The residual trio problem is ள்/ல் separation, not ழ்.

Every error is a phonetic minimal pair

The low-scoring letters don't fail randomly — they confuse with their nearest articulatory neighbours, exactly the families a phonetician would predict. That points the data effort precisely.

Lateralsள் ↔ ல் — the single tightest confusion
Nasalsன் · ண் · ங் · ந் — confuse among themselves
Rhoticsர் ↔ ற் — Tamil's two r-sounds
Short / long vowelsஉ↔ஊ · ஒ↔ஓ — back-vowel duration pairs
Glideய் → இ / ஈ — the glide sounds like its target vowel
06

Single model vs ensemble

The Tamil-phoneme specialist vs the balanced ensemble — the deployment trade.

MetricTamil-phonemeEnsemble
30-way accuracy56%61%
Consonant mean TPR@10%72%79%
Vowel mean TPR@10%91%94%
ழ் @10%75%67%
ள் @10%44%58%
ல் @10%42%58%

The phoneme model wins ழ் outright; the ensemble wins everywhere else and overall. Deployable choice: the ensemble (which contains the phoneme model, keeping most of its ழ் strength).

07

The speakers behind the data

Every number above comes from these voices. Their make-up is also the clearest case for the corpus.

20speakers in the training pool (2,301 clips)
1confirmed female voice
2children (age 9); ages span 9–45
14speakers with no age/gender label at all
Demographic labelling across the 20 speakers
6 labelled
14 unlabelled
age/gender known — 5 studio speakers + 1 test volunteer recordings, names only, no metadata

Of the 6 labelled speakers, gender is known for only 3 — two male, one female — with the rest recorded as “adult” / “kid”. Ages: two children (9), one youth (19), three adults (32 / 44 / 45). No dialect or region is recorded for anyone.

What this points to

These cross-speaker results are strong despite a training pool with one female voice, no dialect coverage, and 14 of 20 speakers unlabelled. A balanced, labelled corpus — gender, age, and region — is exactly what turns a promising research result into a deployable, equitable Tamil tutor.

08

What's next — the roadmap

The numbers above are today's baseline. These are the goals we're working toward — differentiated by acoustic difficulty, because a single number would hide that vowels are near-solved while the sonorant cluster is the real challenge.

Metric (cross-speaker)NowTarget
Vowels (12)94%≥ 95%
Distinct-place consonants83–98%≥ 92%
Retroflex/alveolar sonorant cluster (ண் ன் ர் ற் ல் ள் ழ்)53–67%≥ 80%
Signature sound ழ்75%≥ 85%
Overall (30 letters)61%≥ 85%
Safe-verdict rate — never marks a wrong sound “correct”≥ 95%
Credible, not aspirational

We deliberately do not claim 95% everywhere — ள்ல் and the nasal set overlap acoustically, and even human listeners confuse some as isolated letters. The ≥80% sonorant-cluster target extrapolates a slope we have already shown (ழ் 8%→75%). The gains are data-limited, not method-limited: the approach works; a balanced, labelled corpus plus fine-tuning closes the gap. Measured on a held-out, demographically-balanced test set, reported per-category.