How accurate is Vaklipi?
Word error rate across 11 Indian languages runs 6–16%, depending on the language. Here is exactly how we measured that, language by language, and where it still falls short.
What word error rate means
Word error rate (WER) is the share of words the transcript gets wrong — substitutions, insertions and deletions, divided by the number of words actually spoken. A 10% WER means roughly one word in ten needs a correction. Below about 15%, a transcript is usually faster to fix than to retype. Above about 25%, most people give up and re-listen to the audio, which defeats the point.
How we tested
- Sample: 20 real WhatsApp voice notes per language — not clean studio audio.
- Conditions: a mix of clean speech, phone calls, and noisy environments, in roughly the proportion our users actually upload.
- Comparison: the same audio run through both Sarvam Saarika v2.5 and Whisper large-v3, so the numbers are directly comparable.
- Reported as a range, not a single number, because per-file variance across recording conditions is large enough that a single average would be misleading.
Be clear about what this is: an in-house test on a sample we chose, not an independent academic benchmark. We publish the method so you can weigh it accordingly — and the honest way to check is to run your own audio through the free tier.
Per-language results
| Language | Engine used | Vaklipi WER | Whisper WER |
|---|---|---|---|
| English English | Groq whisper-large-v3-turbo | 5–8% | 5–8% |
| Kannada ಕನ್ನಡ | Sarvam saarika:v2.5 | 10–14% | 28–40% |
| Hindi हिन्दी | Sarvam saarika:v2.5 | 6–10% | 12–18% |
| Tamil தமிழ் | Sarvam saarika:v2.5 | 8–12% | 25–35% |
| Telugu తెలుగు | Sarvam saarika:v2.5 | 9–13% | 26–36% |
| Malayalam മലയാളം | Sarvam saarika:v2.5 | 10–15% | 30–42% |
| Marathi मराठी | Sarvam saarika:v2.5 | 8–12% | 15–22% |
| Bengali বাংলা | Sarvam saarika:v2.5 | 9–13% | 14–20% |
| Gujarati ગુજરાતી | Sarvam saarika:v2.5 | 10–14% | not measured |
| Punjabi ਪੰਜਾਬੀ | Sarvam saarika:v2.5 | 11–15% | not measured |
| Odia ଓଡ଼ିଆ | Sarvam saarika:v2.5 | 12–16% | not measured |
Blank Whisper cells mean we have not run that comparison yet, not that Whisper performs well there. English routes to Whisper because Whisper is genuinely better at it — Sarvam measured 10–14% on Indian-accented English against Whisper's 5–8%.
Why routing beats one model
Most international tools run one English-first model for every language. That is why they report 25–40% error rates on Kannada, Tamil and Telugu. Vaklipi detects the language from the first 60 seconds using Whisper — which is fast and cheap at language ID — then routes Indic audio to Sarvam and everything else to Whisper. You never pay for a transcription run through the wrong engine. If detection is uncertain, you can force the language from a dropdown at upload.
Where it still struggles
Any accuracy page that only lists strengths is marketing. These are the cases where you should expect to do real editing:
- Sung or chanted content
- Music, bhajans, satsangs and anything sung transcribes poorly on every engine we have tested, ours included. We surface a 'music detected' flag when we spot it rather than quietly returning nonsense.
- Heavy overlapping speech
- When three people talk over each other, diarization helps assign turns but the word error rate still climbs. Board meetings with crosstalk are the hardest realistic case.
- Very short clips
- Auto-detection needs signal. Under about a minute, explicitly choosing the language in the dropdown gives noticeably better output than letting detection guess.
- Domain jargon and proper nouns
- Names, legal citations and technical terms are the most common corrections. This is why the inline editor exists — fix a term once and export.
- Poor recording conditions
- Phone-in-pocket audio, wind, and speakerphone-across-a-room recordings degrade results substantially. Placement matters far more than microphone price.
Check a specific language
Test it on your own audio
30 free minutes, no card. Your recordings are the only benchmark that matters.