Access Nepali Conformer (streaming)

This model is released for research and non-commercial use under CC-BY-NC-4.0. Access is reviewed by hand, so tell us who you are and what you plan to build. Requests are usually answered within a couple of days.

We ask for these details so we know who is building on the model and can reach you about corrections, new releases and benchmark changes. We do not share them.

Log in or Sign Up to review the conditions and access this model content.

nepali-conformer-streaming

Cache-aware streaming Nepali ASR (520 ms lookahead). Carries a large, honestly-reported streaming-lineage penalty on real calls — read RESULTS.md before choosing this over the offline model; it exists because a phone agent needs incremental output.

Try it: demo Space · Everything else: github.com/Ampixa/nepaliconformer (NepTel benchmark, per-system outputs, full honest results)

Numbers (measured, not marketed)

benchmark WER
NepTel — real Nepali call audio, human-reviewed refs 59.87
Held-out gold read Nepali (W1 slice) 31.5
Whisper-large-v3 zero-shot on the same NepTel audio 96.3

Architecture

121.3M-parameter 17-layer Conformer (d=512, striding ×4, 40 ms frames), hybrid TDT/CTC decoder, 1,024-piece Devanagari SentencePiece. Chunked-limited attention [[70,13],[70,6],[70,1],[70,0]], fully causal convolutions, cache-aware incremental decoding.

Training data

~1,655 h of mostly conversational Nepali (YouTube podcasts/interviews) with Google Chirp 2 pseudo-labels + 105 h human-labeled read speech; telephony codec, noise, reverb and tempo augmentation. Label-noise ceiling and every measured limitation (English, sung speech, slow speech, end-of-turn) are documented in the repo's RESULTS.md.

Usage

from nemo.collections.asr.models import EncDecHybridRNNTCTCBPEModel
m = EncDecHybridRNNTCTCBPEModel.restore_from("nepali_conformer_streaming.nemo")
print(m.transcribe(["audio.wav"])[0].text)

License: CC-BY-NC-4.0 (weights). Code in the repo: MIT.

Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results

  • Real-call WER on NepTel v0.1 (real Nepali call-center audio, human-reviewed)
    self-reported
    59.870
  • Real-call CER on NepTel v0.1 (real Nepali call-center audio, human-reviewed)
    self-reported
    41.080
  • Read-speech WER on Held-out gold read Nepali (W1 read slice, OpenSLR-54 utterances absent from training)
    self-reported
    31.500