⚠ Official Notice: www.ijisrt.com is the official website of the International Journal of Innovative Science and Research Technology (IJISRT) Journal for research paper submission and publication. Please beware of fake or duplicate websites using the IJISRT name.



Problems of Artificial Intelligence for the Azerbaijani Language: The Impact of Limited Language Corpus


Authors : Raksana Aliashrafova; Namiq Abdurahmanov; Gunel Aghammadova; Vafa Atayeva

Volume/Issue : Volume 11 - 2026, Issue 7 - July


Google Scholar : https://tinyurl.com/mtrhwp9p

Scribd : https://tinyurl.com/2mdhrh96

DOI : https://doi.org/10.38124/ijisrt/26jul101

Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.


Abstract : Artificial intelligence (AI) systems, such as large-scale speech recognition systems, are vital for acquiring linguistic competence and achieving this is essential for their use. Thus, for large-scale speech recognition systems, acquisition of linguistic competence is essential for usage of the system. Even for the widely spoken language like English, Mandarin and Spanish, these corpora are vast and have hundreds of billions of tokens, which can be used to power up virtually all natural-language processing (NLP) tasks. In the online world, however, the Azerbaijani language is far less well resourced: publicly available language corpora are limited to a small number of high-resource languages. In the article, the main problems related to corpus scarcity, such as poor quality of machine translation, incorrect morphological processing, recognition dialects of the Azerbaijan language and failure of chatbots to understand the Azerbaijan language are discussed.

Keywords : Azerbaijani Language, Natural Language Processing, Language Corpus, Low-Resource Languages, Machine Translation, Aİ Language Models, Turkic Languages.

References :

  1. Hajili, N. (2024). Open foundation models for the Azerbaijani language. arXiv preprint arXiv:2407.02337.
  2. Jafarov, Y. M., & Gasimova, R. T. (2020). Some problems in the design of domain names with the Azerbaijani alphabet on the internet. Journal of Information Society Problems, 11(1), 45–52.
  3. Kishiyev, H. (2023). Introducing AzCorpus: A game-changing resource for NLP applications in the Azerbaijani language. Medium.
  4. Liu, Y., Ott, M., Goyal, N., et al. (2019). RoBERTa: A robustly optimised BERT pretraining approach. arXiv preprint arXiv:1907.11692.
  5. Nouri, M., Amani, M., Zohrabi, R., & Asgari, E. (2023). The language model, resources, and computational pipelines for the under-resourced Iranian Azerbaijani. Findings of ACL: AACL-IJCNLP 2023.
  6. Salimov, R., & Mammadova, G. (2023). Evaluation of neural machine translation for the Azerbaijani–English language pair. Proceedings of the Workshop on Language Resources for Turkic Languages.
  7. Sorokin, A., Shavrina, T., & Lyashevskaya, O. (2021). A large-scale study of machine translation in the Turkic languages. Proceedings of ACL 2021.
  8. Karimova, S., & Pecina, P. (2023). Integrated approach to adapting open-source AI models for machine translation of low-resource Turkic languages. Computers, 15(2), 73.
  9. Kartal, J., et al. (2024). Enhancing language learning through technology: A new English–Azerbaijani parallel corpus. arXiv preprint arXiv:2407.05189.

Artificial intelligence (AI) systems, such as large-scale speech recognition systems, are vital for acquiring linguistic competence and achieving this is essential for their use. Thus, for large-scale speech recognition systems, acquisition of linguistic competence is essential for usage of the system. Even for the widely spoken language like English, Mandarin and Spanish, these corpora are vast and have hundreds of billions of tokens, which can be used to power up virtually all natural-language processing (NLP) tasks. In the online world, however, the Azerbaijani language is far less well resourced: publicly available language corpora are limited to a small number of high-resource languages. In the article, the main problems related to corpus scarcity, such as poor quality of machine translation, incorrect morphological processing, recognition dialects of the Azerbaijan language and failure of chatbots to understand the Azerbaijan language are discussed.

Keywords : Azerbaijani Language, Natural Language Processing, Language Corpus, Low-Resource Languages, Machine Translation, Aİ Language Models, Turkic Languages.

Paper Submission Last Date
31 - July - 2026

SUBMIT YOUR PAPER CALL FOR PAPERS
Video Explanation for Published paper

Never miss an update from Papermashup

Get notified about the latest tutorials and downloads.

Subscribe by Email

Get alerts directly into your inbox after each post and stay updated.
Subscribe
OR

Subscribe by RSS

Add our RSS to your feedreader to get regular updates from us.
Subscribe