Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

Dataset Curation for Kalabari NMT System

Oliseamaka T. Olise, Prof V.I.E. Anireh, E.O. Bennett, PhD, Prof. N. Nwiabu

Abstract

The development of Natural Language Processing tools for endangered and low- resource languages is fundamentally hindered by the scarcity of high-quality parallel data. Data for languages like Kalabari is not only scarce but often noisy, inconsistently digitized, and orthographically unstandardized. While prior work has leveraged religious texts for corpus creation, the specific challenges of extracting and normalizing morphologically rich languages with complex diacritics remain underexplored. This paper addresses this gap by introducing a reproducible, modular curation methodology tailored for such languages. We document a six-step pipeline that transforms raw digital texts—sourced from the Kalabari Bible (FiaFia Biabulu) and instructional literature (Kalabari Lingua)—into a clean, verse- aligned, 10,222-pair parallel corpus. We demonstrate that enforcing Normalization Form C is critical for preserving sub-dot diacritics, and we validate the corpus by training a baseline Transformer NMT system. Our contributions are threefold: (1) a transferable curation framework for endangered languages, (2) the first sizable Kalabari-English parallel corpus, and (3) baseline experiments that reveal both the promise and the hallucination pitfalls of training on highly constrained, domain-specific data.

Keywords

Natural Language ProcessingCorpus CurationLow-Resource LanguagesKalabariParallel DataMachine TranslationUnicode Normalization.

References

Ahmadi, S. (2020). Building a corpus for the Zaza–Gorani language family. In Proceedings of the 7th Workshop on NLP for Similar Languages, Varieties and Dialects (pp. 70–78). International Committee on Computational Linguistics. https://aclanthology.org/2020.vardial-1.7/ Ahmadi, S., Azin, Z., Belelli, S., & Anastasopoulos, A. (2023). Approaches to corpus creation for low-resource language technology: The case of Southern Kurdish and Laki. In Proceedings of the Second Workshop on NLP Applications to Field Linguistics (pp. 52– 63). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.fieldmatters-1.7 Artemova, E., Burchell, L., Dementieva, D., Okabe, S., Shmatova, M., & Suarez, P. O. (2025). Low-resource, high-impact: Building corpora for inclusive language technologies. arXiv preprint arXiv:2512.14576. https://doi.org/10.48550/arXiv.2512.14576 Burda-Lassen, O. (2022, September). Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages. In Proceedings of the 15th Biennial Conference of the Association for Machine Translation in the Americas (Workshop 2: Corpus Generation and Corpus Augmentation for Machine Translation) (pp. 28–31). Christodouloupoulos, C., & Steedman, M. (2015). A massively parallel corpus: The Bible in 100 languages. Language Resources and Evaluation, 49(2), 375–395. https://doi.org/10.1007/s10579-014-9287-y Ghimire, R. R., Subedi, B., Prasain, B., Poudyal, P., Acharya, P., Karki, N., ... & Bal, B. K. (2026). NepTam: A Nepali-Tamang parallel corpus and baseline machine translation experiments. arXiv preprint arXiv:2603.14053. https://arxiv.org/abs/2603.14053 Grabar, N., Kanishcheva, O., & Hamon, T. (2018). Multilingual aligned corpus with Ukrainian as the target language. SLAVICORP 2018, Prague, Czech Republic. https://shs.hal.science/halshs-01968343/document Hasan, T., Bhattacharjee, A., Samin, K., Hasan, M., Basak, M., Rahman, M. S., & Shahriyar, R. (2020, November). Not low-resource anymore: Aligner ensembling, batch filtering, and new datasets for Bengali-English machine translation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 2612–2623). Association for Computational Linguistics. Koc, V. (2025). Generative AI and large language models in language preservation: Opportunities and challenges. arXiv preprint arXiv:2501.11496. https://arxiv.org/abs/2501.11496 Mahfuz, T., Dey, S. K., Naswan, R., Adil, H., Sayeed, K. S., & Shahgir, H. S. (2025, January). Too late to train, too early to use? A study on necessity and viability of low-resource Bengali LLMs. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 1183–1200). https://aclanthology.org/2025.coling-main.79/ Marashian, A., Rice, E., Gessler, L., Palmer, A., & von der Wense, K. (2025, January). From priest to doctor: Domain adaptation for low-resource neural machine translation. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 7087–7098). https://aclanthology.org/2025.coling-main.472/ Ndimbo, E. V., Luo, Q., Fernando, G. C., Yang, X., & Wang, B. (2025). Leveraging retrieval- augmented generation for Swahili language conversation systems. Applied Sciences, 15(2), 524. https://www.mdpi.com/2076-3417/15/2/524 Nigatu, H. H., Tonja, A. L., Rosman, B., Solorio, T., & Choudhury, M. (2024, November). The Zeno’s paradox of ‘low-resource’ languages. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 17753–17774). https://aclanthology.org/2024.emnlp-main.983/ Okabe, S., Hämmerl, K., & Fraser, A. (2025, July). Improving parallel sentence mining for low-resource and endangered languages. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 196–205). https://aclanthology.org/2025.acl-short.17/ Olise, O. T., Anireh, V. I. E., Bennett, E. O., & Nwiabu, N. (in press). Democratizing machine translation: A CPU-centric training pipeline for low-resource languages (A Kalabari case study). International Journal of Computer Science and Mathematical Theory. Rajab, J., Aremu, A., Chimoto, E. A., Dunbar, D., Morrissey, G., Thior, F., ... & Rosman, B. (2025, July). The Esethu framework: Reimagining sustainable dataset governance and curation for low-resource languages. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 30763– 30776). https://doi.org/10.48550/arXiv.2502.15916 Ralethe, S., & Buys, J. (2025, January). Cross-lingual knowledge projection and knowledge enhancement for zero-shot question answering in low-resource languages. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 10111–10124). https://aclanthology.org/2025.coling-main.675/ Schwenk, H., Chaudhary, V., Sun, S., Gong, H., & Guzmán, F. (2021). WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 1351–1361). https://aclanthology.org/2021.eacl- main.115/ Shandilya, B., Buchholz, M., & Palmer, A. (2026). GlossAssist—A tool to simplify corpus creation and study the effect of NLP models in low-resource documentation settings. arXiv preprint arXiv:2606.04367. https://arxiv.org/pdf/2606.04367 Yıldız, E., Tantuğ, A., & Diri, B. (2014). The effect of parallel corpus quality vs size in English- to-Turkish SMT. Computer Science & Information Technology, 4, 21–30. https://www.researchgate.net/publication/269162760

More Articles from WORLD JOURNAL OF INNOVATION AND MODERN TECHNOLOGY