Menu
Close

Teaching machines to listen — the quiet revolution in Nepali NLP

Low-resource languages have long been overlooked by mainstream AI. Inside the GCST lab building the first open Nepali speech corpus.

  • calendar
    Aug 6, 2026
  • clock
    5 min read
Teaching machines to listen — the quiet revolution in Nepali NLP

For a decade, natural language processing has been a story told mostly in English. The models that translate, transcribe, and reason are trained on text corpora scraped from a web that overwhelmingly speaks the language of Silicon Valley. Nepali — spoken by more than 30 million people — sits on the periphery of that conversation.

A corpus built by hand

At GCST, a team of undergraduates and faculty are building what we believe will be the largest openly-licensed Nepali speech corpus in existence. It is slow, unglamorous work: recording, transcribing, aligning, and validating hundreds of hours of audio across dialects.

“We are not trying to beat OpenAI. We are trying to make sure the future speaks our language.”

Why open matters.

The dataset will be released under a permissive license. Any researcher, startup, or student anywhere in the world will be able to build on top of it — from voice assistants to accessibility tools for low-literacy users. The initial release includes:

  • 12,000 hours of raw audio
  • Aligned transcripts, verified by native speakers
  • Baseline ASR and TTS models fine-tuned on Whisper and VITS
  • Full dataset documentation, provenance, and consent records

The first public release drops in autumn 2026. If you are a researcher who wants early access, we would love to hear from you.