From generative AI to autonomous systems, discover the technologies reshaping industries and everyday.
Low-resource languages have long been overlooked by mainstream AI. Inside the GCST lab building the first open Nepali speech corpus.
For a decade, natural language processing has been a story told mostly in English. The models that translate, transcribe, and reason are trained on text corpora scraped from a web that overwhelmingly speaks the language of Silicon Valley. Nepali — spoken by more than 30 million people — sits on the periphery of that conversation.
At GCST, a team of undergraduates and faculty are building what we believe will be the largest openly-licensed Nepali speech corpus in existence. It is slow, unglamorous work: recording, transcribing, aligning, and validating hundreds of hours of audio across dialects.
“We are not trying to beat OpenAI. We are trying to make sure the future speaks our language.”
The dataset will be released under a permissive license. Any researcher, startup, or student anywhere in the world will be able to build on top of it — from voice assistants to accessibility tools for low-literacy users. The initial release includes:
The first public release drops in autumn 2026. If you are a researcher who wants early access, we would love to hear from you.