How we got here — the model ladder
Recognition of ழ், Tamil's hardest sound, at each step. Every rung is a more Tamil-aware model.
(formants + CNN)
(128-lang)
IndicWav2Vec
(sound-trained)
(all four)
ழ் recognition climbs 8% → 42% → 53% → 75% as models get more Tamil-aware — but no single model wins everything. The system we built is an ensemble that fuses AI4Bharat's IndicWav2Vec with three other pretrained encoders; it trades a little peak accuracy for balance across all 30 letters.
Consonants — மெய் (18)
Cross-speaker recognition (TPR) at a fixed 10% false-accept budget. The small grey number is 30-way cross-letter accuracy — how often the letter beats all 29 others outright.
Vowels — உயிர் (12)
Tamil vowels are acoustically distinct — the ensemble separates them almost perfectly.
What's solved, what's left
Vowels (94%) and distinct-place consonants — ச் க் ஞ் த் ப் ம் ட் all 83–98%. The pretrained ensemble handles the bulk of the alphabet cross-speaker.
The retroflex & alveolar sonorants — ண் ன் ர் ற் ல் ள் ழ் — 53–78%. The tightest single confusion is ள் ↔ ல், the two laterals. This is exactly where more speaker-diverse data or cluster-specific fine-tuning pays off.
The errors, decoded
Where the residual confusion actually lands — and why it's phonetically principled, not random.
| true ↓ / pred → | ழ் | ள் | ல் | recall |
|---|---|---|---|---|
| ழ் | 23 | 10 | 3 | 64% |
| ள் | 11 | 11 | 14 | 31% |
| ல் | 7 | 17 | 12 | 33% |
ழ் is the strong one (64%). The real mush is ள் ↔ ல் — the two laterals bleeding into each other almost symmetrically (14 and 17). The residual trio problem is ள்/ல் separation, not ழ்.
The low-scoring letters don't fail randomly — they confuse with their nearest articulatory neighbours, exactly the families a phonetician would predict. That points the data effort precisely.
Single model vs ensemble
The Tamil-phoneme specialist vs the balanced ensemble — the deployment trade.
| Metric | Tamil-phoneme | Ensemble |
|---|---|---|
| 30-way accuracy | 56% | 61% |
| Consonant mean TPR@10% | 72% | 79% |
| Vowel mean TPR@10% | 91% | 94% |
| ழ் @10% | 75% | 67% |
| ள் @10% | 44% | 58% |
| ல் @10% | 42% | 58% |
The phoneme model wins ழ் outright; the ensemble wins everywhere else and overall. Deployable choice: the ensemble (which contains the phoneme model, keeping most of its ழ் strength).
The speakers behind the data
Every number above comes from these voices. Their make-up is also the clearest case for the corpus.
Of the 6 labelled speakers, gender is known for only 3 — two male, one female — with the rest recorded as “adult” / “kid”. Ages: two children (9), one youth (19), three adults (32 / 44 / 45). No dialect or region is recorded for anyone.
These cross-speaker results are strong despite a training pool with one female voice, no dialect coverage, and 14 of 20 speakers unlabelled. A balanced, labelled corpus — gender, age, and region — is exactly what turns a promising research result into a deployable, equitable Tamil tutor.
What's next — the roadmap
The numbers above are today's baseline. These are the goals we're working toward — differentiated by acoustic difficulty, because a single number would hide that vowels are near-solved while the sonorant cluster is the real challenge.
| Metric (cross-speaker) | Now | Target |
|---|---|---|
| Vowels (12) | 94% | ≥ 95% |
| Distinct-place consonants | 83–98% | ≥ 92% |
| Retroflex/alveolar sonorant cluster (ண் ன் ர் ற் ல் ள் ழ்) | 53–67% | ≥ 80% |
| Signature sound ழ் | 75% | ≥ 85% |
| Overall (30 letters) | 61% | ≥ 85% |
| Safe-verdict rate — never marks a wrong sound “correct” | — | ≥ 95% |
We deliberately do not claim 95% everywhere — ள்↔ல் and the nasal set overlap acoustically, and even human listeners confuse some as isolated letters. The ≥80% sonorant-cluster target extrapolates a slope we have already shown (ழ் 8%→75%). The gains are data-limited, not method-limited: the approach works; a balanced, labelled corpus plus fine-tuning closes the gap. Measured on a held-out, demographically-balanced test set, reported per-category.