01/07/2026
Releasing a 500,000 Parallel Corpus for South Sudanese Languages
I'm excited to announce the upcoming release of a 500,000-sentence parallel corpus covering English – Nuer – Dinka on Hugging Face.
This dataset is being made freely available to support the global AI and NLP community. My goal is to help researchers, students, and developers build machine translation systems, multilingual language models, chatbots, speech technologies, and other AI applications for South Sudanese languages.
By making this resource open, I hope to reduce the barriers to developing high-quality language technologies for communities that have long been underrepresented in AI.
Whether you're training a translation model, fine-tuning an LLM, or conducting NLP research, this dataset provides a strong foundation for your work.
Dataset: https://huggingface.co/datasets/Tajnaam/english-nuer-dinka-parallel-corpus
I look forward to seeing what the community builds with it. Contributions, feedback, and collaborations are always welcome.