References
Ahmadi, S. (2020). Building a corpus for the Zaza–Gorani language family. In Proceedings of the 7th Workshop on NLP for Similar Languages, Varieties and Dialects (pp. 70–78). International Committee on Computational Linguistics. https://aclanthology.org/2020.vardial-1.7/ Ahmadi, S., Azin, Z., Belelli, S., & Anastasopoulos, A. (2023). Approaches to corpus creation for low-resource language technology: The case of Southern Kurdish and Laki. In Proceedings of the Second Workshop on NLP Applications to Field Linguistics (pp. 52– 63). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.fieldmatters-1.7 Artemova, E., Burchell, L., Dementieva, D., Okabe, S., Shmatova, M., & Suarez, P. O. (2025). Low-resource, high-impact: Building corpora for inclusive language technologies. arXiv preprint arXiv:2512.14576. https://doi.org/10.48550/arXiv.2512.14576 Burda-Lassen, O. (2022, September). Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages. In Proceedings of the 15th Biennial Conference of the Association for Machine Translation in the Americas (Workshop 2: Corpus Generation and Corpus Augmentation for Machine Translation) (pp. 28–31). Christodouloupoulos, C., & Steedman, M. (2015). A massively parallel corpus: The Bible in 100 languages. Language Resources and Evaluation, 49(2), 375–395. https://doi.org/10.1007/s10579-014-9287-y Ghimire, R. R., Subedi, B., Prasain, B., Poudyal, P., Acharya, P., Karki, N., ... & Bal, B. K. (2026). NepTam: A Nepali-Tamang parallel corpus and baseline machine translation experiments. arXiv preprint arXiv:2603.14053. https://arxiv.org/abs/2603.14053 Grabar, N., Kanishcheva, O., & Hamon, T. (2018). Multilingual aligned corpus with Ukrainian as the target language. SLAVICORP 2018, Prague, Czech Republic. https://shs.hal.science/halshs-01968343/document Hasan, T., Bhattacharjee, A., Samin, K., Hasan, M., Basak, M., Rahman, M. S., & Shahriyar, R. (2020, November). Not low-resource anymore: Aligner ensembling, batch filtering, and new datasets for Bengali-English machine translation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 2612–2623). Association for Computational Linguistics. Koc, V. (2025). Generative AI and large language models in language preservation: Opportunities and challenges. arXiv preprint arXiv:2501.11496. https://arxiv.org/abs/2501.11496 Mahfuz, T., Dey, S. K., Naswan, R., Adil, H., Sayeed, K. S., & Shahgir, H. S. (2025, January). Too late to train, too early to use? A study on necessity and viability of low-resource Bengali LLMs. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 1183–1200). https://aclanthology.org/2025.coling-main.79/ Marashian, A., Rice, E., Gessler, L., Palmer, A., & von der Wense, K. (2025, January). From priest to doctor: Domain adaptation for low-resource neural machine translation. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 7087–7098). https://aclanthology.org/2025.coling-main.472/ Ndimbo, E. V., Luo, Q., Fernando, G. C., Yang, X., & Wang, B. (2025). Leveraging retrieval- augmented generation for Swahili language conversation systems. Applied Sciences, 15(2), 524. https://www.mdpi.com/2076-3417/15/2/524 Nigatu, H. H., Tonja, A. L., Rosman, B., Solorio, T., & Choudhury, M. (2024, November). The Zeno’s paradox of ‘low-resource’ languages. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 17753–17774). https://aclanthology.org/2024.emnlp-main.983/ Okabe, S., Hämmerl, K., & Fraser, A. (2025, July). Improving parallel sentence mining for low-resource and endangered languages. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 196–205). https://aclanthology.org/2025.acl-short.17/ Olise, O. T., Anireh, V. I. E., Bennett, E. O., & Nwiabu, N. (in press). Democratizing machine translation: A CPU-centric training pipeline for low-resource languages (A Kalabari case study). International Journal of Computer Science and Mathematical Theory. Rajab, J., Aremu, A., Chimoto, E. A., Dunbar, D., Morrissey, G., Thior, F., ... & Rosman, B. (2025, July). The Esethu framework: Reimagining sustainable dataset governance and curation for low-resource languages. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 30763– 30776). https://doi.org/10.48550/arXiv.2502.15916 Ralethe, S., & Buys, J. (2025, January). Cross-lingual knowledge projection and knowledge enhancement for zero-shot question answering in low-resource languages. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 10111–10124). https://aclanthology.org/2025.coling-main.675/ Schwenk, H., Chaudhary, V., Sun, S., Gong, H., & Guzmán, F. (2021). WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 1351–1361). https://aclanthology.org/2021.eacl- main.115/ Shandilya, B., Buchholz, M., & Palmer, A. (2026). GlossAssist—A tool to simplify corpus creation and study the effect of NLP models in low-resource documentation settings. arXiv preprint arXiv:2606.04367. https://arxiv.org/pdf/2606.04367 Yıldız, E., Tantuğ, A., & Diri, B. (2014). The effect of parallel corpus quality vs size in English- to-Turkish SMT. Computer Science & Information Technology, 4, 21–30. https://www.researchgate.net/publication/269162760