Digital Umuganda Releases First Open-Source African Language AI Datasets for Afrivoice V2

AI Quick Summary
Digital Umuganda has announced the release of open-source datasets for four African languages as part of its Afrivoice V2 initiative. This project aims to develop 10,000 hours of open-source speech data across 20 African languages over the next 12 months, addressing a critical barrier for innovators building AI solutions across the continent.
The initial datasets cover Kirundi, Ndau, Ndebele, and Oshiwambo, serving communities in East, Central, and Southern Africa. This effort seeks to ensure that African languages are adequately represented in future AI technologies.
What is Afrivoice V2?
The Afrivoice V2 initiative focuses on collecting and publishing Automatic Speech Recognition (ASR) datasets. Its goal is to provide 500 hours of transcribed speech for each of 20 additional African languages, totaling 10,000 hours of data. This effort builds on previous work in developing speech data for African languages, including Digital Umuganda's own contributions for Kinyarwanda and Swahili.
The project emphasizes collaboration with local researchers and principal investigators. These individuals are responsible for leading the work on the ground, ensuring an understanding of the linguistic, cultural, and social nuances within their communities.
Initial Datasets Released
As of August 31, 2026, a subset of the first datasets from Afrivoice V2 has been released. These cover Kirundi, Ndau, Ndebele, and Oshiwambo. The release marks a step towards expanding Africa's language resources for AI development, making these languages more represented in the global digital landscape.
Afrivoice V2 at a glance:
Goal: 10,000 hours of open-source speech data.
Languages: 20 African languages.
Per language: 500 hours of transcribed speech.
Initial release: Kirundi, Ndau, Ndebele, Oshiwambo.
Why it Matters
Digital Umuganda states that the future of Africa's AI-enabled digital economy requires AI models that understand local languages and contexts. This initiative aims to make AI truly inclusive by providing foundational data for developers. These datasets can support the creation of locally relevant technologies, such as voice-enabled health information services, educational platforms, and farmer advisory systems.
George Mandhlazi, Project Manager for the Ndau language in Zimbabwe, highlighted the importance of digital visibility. He stated that the project helps ensure Ndau speakers can access essential services in their own language. Hee-Dee Walenga, Oshiwambo Project Manager in Namibia, noted the initiative's role in democratizing technology access, preventing exclusion based on language.
How to Access
Researchers and developers can access the released datasets through the Afrivoice_V2 Hugging Face repository. Digital Umuganda plans to release additional datasets in the coming months.
If you enjoyed this article, follow us on WhatsApp for daily tech updates. If you have an idea, need to be featured or need to partner, reach out to us at editorial@techinika.com or use our contact page.
Don't let the story end here.
Share your thoughts, ask questions, and connect with the community.

Cishahayo Songa Achille
Chief EditorCishahayo Songa Achille is a Rwandan software engineer and tech entrepreneur focused on democratizing digital skills. He is best known as the Founder and Managing Director of Techinika, an edtech firm established in 2020 to make complex technological advances accessible to the general public and build solutions for the biggest problems.
View all articles by Cishahayo Songa Achille →Up Next
AI Is Speeding Up Open Source Security, But Developers Still Have the Final SayBy ISHIMWE Jean Claude • 3 min read


