Dia Krishnan, Kavya Manohar, Kumarmanas Nethil
Guide
15
min read
Diagnostic ASR evaluation for Indic and domain-specific speech
SCRIBE, the open-source evaluation tool that steered our system to the DISPLACE 2026 leaderboard, is now on PyPI.
TL;DR — A machine transcript is useful only when fixing it is faster than typing it fresh, and that depends on which mistakes it makes, not how many. The standard accuracy score treats a missing comma and a wrong date as the same error — and it structurally over-penalises many Indian languages, where the same speech can legitimately be written as one compound word or two. SCRIBE scores words, numerals, punctuation and your own domain terms separately, so the score tells you what to fix. On the DISPLACE 2026 medical-speech challenge, it showed our "medical vocabulary problem" didn't exist — the real losses were elsewhere — and led us to spend the remaining weeks on the fixes that actually moved the score. pip install scribe-eval· Paper at Interspeech 2026.
At Adalat AI, we build speech recognition for Indian courtrooms. The test our work faces every day is a stenographer's: is correcting this transcript faster than typing it fresh? For an automatic speech recognition (ASR) system to actually be useful, its output needs the right punctuation, numerals in the required format — often more than one format is equally right — and the spellings the courtroom expects.
Whether a draft passes that test depends on the kind of its errors, not the count. A missing comma costs a keystroke. A misheard medicine name or a scrambled hearing date sends the typist back to the audio — and after a few such cases, the stenographer's trust in the system is lost.
The number nearly every ASR system is judged by is called Word Error Rate (WER). It aligns the system's transcript against a human reference and counts substitutions, deletions and insertions, divided by the number of reference words. Here is a Malayalam sentence, and a transcript of it, scored with WER:
Reference: ഇങ്ങനെ ഉള്ള ഒരു കഥാപാത്രമായിട്ടാണ് ചിത്രീകരിക്കുന്നത്
Hypothesis: ഇങ്ങനെയുള്ള ഒരു കഥാപാത്രം ആയിട്ടാണ് ചിത്രീകരിക്കുന്നത്.
WER: 100%

Every word is counted wrong. But every word is right. The transcript joined ഇങ്ങനെ ഉള്ള into one word, split കഥാപാത്രമായിട്ടാണ് into two, and added a full stop. A person editing this draft would change nothing.
This is not a rare failure mode. In Dravidian languages — Malayalam, Kannada, Tamil, Telugu — words agglutinate: the same speech can be legitimately written as one word or two, and the spelling changes at the junction when words fuse (notice the യ that appeared inside ഇങ്ങനെയുള്ള — stripping spaces does not make the strings equal). WER aligns words one-to-one, so a merge gets counted twice - once as a wrong word, once as a missing word; in our measurements this alone inflates error rates by up to 30% relative (the SCRIBE paper has the details). It is a structural penalty against an entire language family.
The second failure is subtler, and not limited to Dravidian languages: WER weighs every error the same. Here is a Hindi example:
Reference: भाई साहब ने 250 रुपये दिए थे
Hypothesis: भाईसाहब ने 205 रुपये दिए थे
WER: 43% — three errors: भाई → भाईसाहब, साहब dropped, 250 → 205

Two of those three "errors" are one spacing convention — भाई साहब written solid; the third changed an amount of money. For the score they are interchangeable; for the person checking the record they are nothing alike. Acoustic failures, numeral formatting and punctuation, all collapsed into one scalar that tells you nothing about what or where to fix.
So the first thing an evaluation should do is stop adding unlike things into one number — or rather, add them only after telling them apart. Split the transcript into three buckets: words (lexical tokens), numbers, and punctuation. Score each bucket separately, and treat a sandhi merge or split as one boundary mistake, not two. The category rates then add back into one headline figure, with every point of it accounted for. On the Hindi sentence above that reads: lexical 0% (the compound recognised as a sandhi match), numeral 14%, punctuation 0% — a total of 14%, against WER's 43%. And each rate points at a different fix: the lexical rate at the acoustic model, the numeral rate at formatting, the punctuation rate at post-processing.
This is SCRIBE, our diagnostic evaluation framework for ASR — named for the role it measures: whether a system can serve as a reliable scribe. Its headline figure, WER_SCRIBE, is the sum in the cards above. It is accepted at Interspeech 2026, and it is now on PyPI:
SCRIBE ships with these three categories built in, and any category can be added or modified: give it a plain-text list of terms — or regex patterns — and it scores that class separately from the rest of the text, alongside tables of the most frequent substitutions, deletions, insertions and sandhi mismatches. In our courtrooms, that means act names, section numbers and exhibit references tagged LEGAL, with their own error rate; in a clinic it is drug names and dosages. Where the model is struggling is the signal a developer needs to decide the development strategy. The rest of this post is what that looked like in practice, at the DISPLACE 2026 challenge.
DISPLACE 2026: the challenge and the metrics
DISPLACE 2026 (DISPLACE-M) — DIarization of SPeaker and LAnguage in Conversational Environments — is organised by IISc Bangalore with support from NLTM–MeitY–Bhashini. It is a challenge series built around several people talking naturally, switching languages mid-sentence, recorded far-field on a single microphone in real rooms. The 2026 edition moves that setting into healthcare: community health workers (ASHAs) in conversation with patients, in multilingual, code-mixed Hindi and Kannada.
The setting is different from ours; the shape is the same. Our courtrooms have many people talking — judge, counsel, witnesses, clerks — often over each other, in more than one language, in rooms not designed for recording. Diarization and speaker-attributed ASR on code-mixed, far-field, overlapping speech is the problem we care about deeply. DISPLACE-M defined the same problem, with ASHAs and patients in place of judges and counsel.
In both places, a transcription error is not a benchmark statistic — it lands on people and their life. That is why we care not just about how much our systems get wrong, but about what kind of wrong it is.
Adalat AI currently sits at the top of the leaderboard in both the Automatic Speech Recognition and Speaker Diarization tracks of DISPLACE-M. Those systems are the subject of a separate technical report; this post is about the evaluation work that ran alongside them — the analysis that decided where our modelling effort went.


A quick orientation on how the two tracks are scored. Speaker diarization answers who spoke when: the system takes a long recording and returns time segments labelled by speaker. It is scored with Diarization Error Rate (DER) — the fraction of audio time that is attributed to the wrong speaker, missed, or falsely detected as speech. Lower is better.
ASR answers what was said. In a multi-speaker conversation, a transcript is only useful if the words are also attached to the right speaker at the right time — so DISPLACE scores ASR with tcpWER (time-constrained minimum-permutation WER, from the MeetEval toolkit). Informally: WER, but a word only counts as correct if it is placed with the right speaker in roughly the right place in time. Again, lower is better.
The rest of this post is about the ASR side: what tcpWER couldn't tell us, what SCRIBE could, and the many things we discovered on the way about speech in the wild — code-mixed, colloquial speech with domain-relevant terminology.
What we built
Our ASR systems pair a Fast Conformer encoder with an RNN-T (transducer) decoder: a fine-tuned IndicConformer for Hindi, and a pretrained multilingual model for Kannada, decoded greedily for the diagnostic runs described below. A technique that comes up below is shallow fusion: interpolating an external n-gram language model's scores with the acoustic model's during decoding, to bias it toward in-domain word sequences.
The full story of how these systems were built — exact architectures, the checkpoints each model was fine-tuned from, additional datasets, language-ID prompting, and the shallow-fusion LM setup — will appear in a separate technical report on our DISPLACE-M system development. The diarization system is its own line of work, covered in the same report; SCRIBE played no part in it — it is a transcript-level tool, and who spoke when is outside what it measures. This post stays with the evaluation side.
The plateau we couldn't explain
DISPLACE-M released train and validation data with transcripts. Our fine-tuned models had plateaued on the validation data. Since this is a medical corpus, we assumed the models needed better medical-domain language modelling. So we tried the standard techniques: an n-gram language model over the domain text for shallow fusion, and word boosting with a curated list of medical terms. Neither gave the expected results.
Boosting made Hindi slightly worse at every weight.
A Kannada LM built from formal medical text made WER worse by 17% (absolute); a much smaller LM built only from the conversational transcripts improved it by 2.6% (absolute).
The challenge rules allowed heavier options — external data, LLM post-correction, which the Phase-1 winner had used — but before spending effort there, we wanted to know why the medical-domain work wasn't paying off. This is exactly the where is it struggling question, so we pointed SCRIBE at it.
What SCRIBE told us
We extracted a plausible list of medical terms from the dev-set transcripts — covering drugs, symptoms, anatomy, procedures, maternal and paediatric vocabulary, and health-system terms, with units and dosages as regex patterns — audited it by hand and with LLMs, and packaged it as a SCRIBE config per language:
Language | Literal terms | Regex patterns | Evaluation set |
|---|---|---|---|
Hindi | 233 | 9 | 24 recordings, 5,620 segments |
Kannada | 163 | 6 | 5 recordings, 577 segments |
We then ran SCRIBE over held-out validation hypotheses from the greedy-decoded RNN-T systems we had at the time.
A note on terminology before the numbers: SCRIBE counts tokens. Every regular word separated by spaces is a SCRIBE token; punctuation marks and standalone numerals are their own tokens, and a multi-word medical term defined as a single config entry (say, आयरन की गोली) counts as one token. We use "token" consistently below.
1. Medical vocabulary was the easy part
Medical terms are the best-recognised part of the transcript: in Hindi the models get 94% of medical tokens right against 85% of general tokens, and in Kannada 84% against 62%.


Medical terms contribute 0.4 points of Hindi's 16.5% WER and 1.1 points of Kannada's 39.4%. A perfect medical lexicon — every term right, every time — would have moved Hindi by less than half a point.
So where was the error budget actually going? Three things, all visible in SCRIBE's error tables.
2. Where were we losing? Mostly filler words
Fillers were not a category we had defined — our config only marked medical terms. But alongside the per-category rates, SCRIBE's report lists the most frequent substitutions, deletions and insertions over the whole transcript, and those tables made the pattern visible: the words the models lose most often are the small ones.
The top Hindi deletions are है, अच्छा, हां, नहीं, हम्म, हैं, हाँ (hai, achha, haan, nahin, hmm, hain, haan) — one-to-three-phoneme agreement particles. Particles and fillers account for 41% of all Hindi deletions. The same words also sit at the top of the insertion list (है hai 74, तो to 55, नहीं nahin 31 in Hindi).
Kannada is starker: 58% of deletions are fillers, and the hesitation ಅ (a) alone is dropped 113 times — two points of WER on its own. Kannada's dropped hesitations do not reappear anywhere — 429 deletions against only 137 insertions in the whole set. Both patterns are visible in the error tables in the next two sections.
3. Zooming into the Hindi errors
SCRIBE's substitution table for Hindi is topped by pairs that are the same word written two ways: हूँ↔हूं, हाँ↔हां, आँखों↔आंखों, चीज़↔चीज, थोड़ा↔थोड़ा (hoon, haan, aankhon, cheez, thoda — each pair reads identically aloud). The two थोड़ा (thoda) are different byte sequences: the reference uses the precomposed character ड़ — a single Unicode codepoint — while the hypothesis writes ड followed by a combining nukta, two codepoints. They render identically, read identically, and mean the same thing, yet a scorer that compares raw strings counts every occurrence as a substitution. This is a clear case for text normalisation in evaluation.

Above even the variants sits a grammar pair: हैं↔है (hain↔hai, plural versus singular — 96 substitutions one way, 72 the other). In fast conversational speech the two are often genuinely unidentifiable.
Counting over SCRIBE's per-token error output, 515 substitutions are pure orthographic variants — chandrabindu versus anusvara, nukta present or absent, or nothing more than a Unicode composition difference. Such errors are worth about one point of WER. They cannot be recovered by any acoustic or language model, because there is nothing to recover; they dissolve under text normalisation.
Word-boundary (sandhi) variation — आप को↔आपको (aap ko↔aapko, "to you"), ले के↔लेके (le ke↔leke, "having taken") — is the other half of this, the same pattern the भाई साहब example opened with. Strictly speaking, these Hindi pairs are spacing conventions rather than phonological sandhi — nothing changes in the sound or the letters at the junction, unlike the Malayalam example that opened this post. SCRIBE uses sandhi as an umbrella term for both, because the scoring failure is identical: a legitimate boundary written two ways, which a word-level metric charges twice. This correction is worth 0.65 points of WER for Hindi, 5.2 for Kannada. This difference comes from the language family: Kannada is more agglutinative than Hindi. SCRIBE's sandhi detection is an orthographic heuristic, not a full linguistic analysis, so these counts are indicative rather than exact.
4. Zooming into the Kannada errors
The Kannada substitution table generated by SCRIBE gave us more insights. The top entries are pairs like ಅಂದ್ರೆ→ಅಂದರೆ (andre→andare, "that is", 48 times), ಮಾಡ್ಬೇಕು for ಮಾಡಬೇಕು (maadbeku for maadabeku, "must do"), ಅವ್ರಿಗೆ for ಅವರಿಗೆ (avrige for avarige, "to them") — the same word, spelled as it is spoken versus as it is written. During conversational speech, speakers contracted or skipped vowels. To size the pattern we ran a short script over SCRIBE's per-token error output, comparing reference and hypothesis after stripping vowel and virama marks: 791 of 1,663 substitutions — fourteen points of WER — have an identical consonant skeleton.
Our first reading was that the model writes formal, literary Kannada even when speakers contracted vowels at times; the most visible pattern (ಅಂದ್ರೆ→ಅಂದರೆ: andre→andare) says that. But the full table is more interesting. The model goes the other way round too — in 221 pairs the hypothesis is the more contracted, spoken-style form and the reference the formal one, against 162 pairs the reverse. There is no single direction. The model and the reference simply do not share a convention for writing spoken Kannada, and disagreements are scored as errors.

Why did this surface in Kannada and not Hindi? Our tables answer this only partly. Spelling-level disagreement between reference and hypothesis is much more common in Kannada: 48% of its substitutions differ only in vowel and virama marks, against about 11% spelling variants in Hindi. One possible explanation, which we offer as a hypothesis and welcome correction from Kannada speakers, is that written Kannada may not have a single settled spelling for many spoken forms. A word like ಅಂದ್ರೆ (andre) can be written in more than one way, and annotators transcribing the same audio may differ. Hindi shows a milder version of the same effect: our tables contain टैबलेट↔टेबलेट, two accepted spellings of tablet.
If this reading is correct, it is not a flaw of DISPLACE-M; it is a property of transcribing real conversational speech, which this benchmark makes visible. There is no agreed answer yet to how spoken Kannada should be written down — both the model and the references move between formal and colloquial renderings of the same spoken forms — and every disagreement about it is currently scored as an ASR error.
What we did with it
SCRIBE explained the results we already had, and set the direction for the time we had left.
Medical-term boosting had hurt because there was nothing to fix. The terms were already the best-recognised part of the transcript, and the in-domain LM already covered them. We dropped the boosting path and did not pursue LLM post-correction, because it would have spent the remaining time on 0.4% (Hindi) to 1.1% (Kannada) of the error budget.
The Kannada LM result now had a reason. The formal medical corpus was the wrong kind of Kannada — written, textbook style; the small conversational corpus matched how people in these recordings actually talk. SCRIBE's substitution table said the same thing about the acoustic model. So the last Kannada work went into matching the spoken language, not expanding the vocabulary: adapting the pretrained encoder on the conversational Kannada data (−11% absolute WER on held-out validation), and fixing segmentation — splitting over-long diarization segments so the model stopped repeating itself on long segments. No medical data was added.
Hindi kept the conversational n-gram LM and otherwise stayed as it was. What remains of its error budget is particles, spelling conventions and word boundaries, and none of those are fixed by more vocabulary.
One lesson from the Hindi tables is worth stating separately, because it is about evaluation rather than modelling. We did not build any normalisation fix during the challenge — submissions are scored with the organisers' evaluation pipeline — but the analysis makes the general point plainly: an evaluation protocol should normalise, on both reference and hypothesis, the orthographic variants that no model can distinguish — chandrabindu/anusvara, nukta, Unicode composition. Otherwise a scorer charges systems for errors that do not exist. Some of this is entirely mechanical: 213 of the 515 Hindi "variant" substitutions are pure Unicode-composition differences, which standard Unicode normalisation (NFC) of both reference and hypothesis folds before any linguistic normalisation is even needed. DISPLACE-M is the rare challenge built on genuinely spontaneous, code-mixed clinical speech, and it is exactly this kind of data that surfaces these questions for the whole field. We are sharing these observations here — and SCRIBE's own normaliser has room to grow on the same points.
Using SCRIBE
SCRIBE has been through several rounds of internal review since the Interspeech paper and is packaged for general use. It starts from a file you almost certainly already have.
1. A predictions file. One utterance per line in JSONL, with one field for the reference and one for the hypothesis — the field names are yours, and any extra fields are ignored.
2. The error report. Point compute_sample_errors at the file and name your two fields
3. Add a domain lens. When you want a term class scored on its own — medical terms here; act names and section numbers in our courtrooms — describe it in a plain-text config of literal terms, regex patterns, or both:
Then hand it to the same call from step 2:
DomainConfig comes from the same scribe import, and bundled starters exist: DomainConfig.legal(), DomainConfig.medical(), DomainConfig.technical().
Every number in this post comes from exactly this workflow on our two validation files.
It is built with Indic orthography in mind, tested on Hindi, Kannada and Malayalam — but nothing in it is Indic-specific. If you have a domain (legal, medical, financial) and a suspicion about where your errors live, SCRIBE will tell you whether you're right before you spend the compute. You can always opt out of Sandhi, if your language does not need that.
PyPI:
pip install scribe-evalCode: github.com/adalat-ai-tech/scribe-eval · archived with a Zenodo DOI
Paper: SCRIBE, Interspeech 2026 · arXiv:2605.20712
Demo: SCRIBE — Diagnostic ASR Evaluation for Indic Languages
Examples: https://github.com/adalat-ai-tech/scribe-eval/tree/main/examples
We'll be presenting SCRIBE at Interspeech 2026 in Sydney, and our DISPLACE-M system report will follow. If you try SCRIBE on your own corpus, we'd love to hear what it tells you.
