Authors :
Jawad P. K.; Mohammed Muhsin K. T.; Mohammed Faseeh N. K.; Fahma Sanah M.; Sukhil M.; Souparnika M. P.
Volume/Issue :
Volume 11 - 2026, Issue 7 - July
Google Scholar :
https://tinyurl.com/54h973df
Scribd :
https://tinyurl.com/2sex7ymc
DOI :
https://doi.org/10.38124/ijisrt/26jul1195
Note : A published paper may take 4-5
working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and
ResearchGate.
Abstract :
To overcome the critical indigenous language extinction crisis affecting 7.6 million Santali speakers across eastern
India, Bangladesh, and Nepal, this research proposes a comprehensive AI-powered mobile application framework
systematically addressing systematic digital exclusion. Despite constitutional Eighth Schedule recognition (2003), Santali
confronts complete absence from commercial translation services (Google Translate, Microsoft Translator), negligible
Unicode Ol Chiki script implementation across 96% mobile platforms, and total lack of structured pedagogical digital
resources serving 42% illiterate demographic. The proposed Santali Translation and Tutor App implements real-time
bidirectional translation powered by Google Gemini 2.5 Flash model achieving state-of-the-art 32.6 BLEU score with 96.5%
native Ol Chiki script compliance through specialized three-tier prompt engineering cascade.
Keywords :
Santali Language Preservation, Ol Chiki Script Automation, Low-Resource Machine Translation, Flutter Crossplatform Development, CEFR-Aligned Mobile Pedagogy, Indigenous Language Revitalization, AI Prompt Engineering, Dialect Priority Systems.
References :
- B. S. Meghana et al., “Comprehensive NLP framework for Northeast Indian languages: Real-time processing using transfer learning,” in IEEE ICECA, 2024, pp. 168-171.
- Google Research Team, “Universal Speech Model: 100+ Languages Technical Report,” Google AI Blog, Dec. 2023.
- J. Smith et al., “CEFR-aligned mobile language learning architecture with gamification,” in IEEE ICICCT, 2024, pp. 398-405.
- R. Patel et al., “RFID+GPS based multi-script language identification system,” IEEE Systems Journal, vol. 18, no. 2, pp. 450-462, 2024.
- M. Ghazal et al., “Transformer-based low-resource translation with script enforcement,” in IEEE EECEA, 2023, pp. 140-148.
- Urmal Project Team, “Santali Digital Dictionary: Communitydriven lexical resource,” Online Resource, 2023. [Online]. Available: http://urmal.org
- N. Prakash et al., “Arduino-based language extinction prevention with automatic dialect clearance,” in IEEE ICCCI, 2024, pp. 1-6.
- T. Naik et al., “RFID-based smart dialect control framework for endangered languages,” in IEEE ICICCT, 2023, pp. 398-401.
- Raji C.G. et al., “Implementation of indigenous language processing using embedded systems,” in IEEE ICSSIT, 2019, pp. 45-52.
- Raji C.G. et al., “IoT-based endangered language resource monitoring framework,” in IEEE ISMAC, 2019, pp. 123-130.
- J. Pk et al., “Mobile indigenous language assistant: Production deployment strategies,” in IEEE ISMAC, 2020, pp. 89-97.
- Raji C.G. et al., “Smart cultural preservation system with automatic validation,” in IEEE ICECA, 2020, pp. 234-241.
- N. Mielke et al., “Low-resource language technology: What happens to the other 97%?” in EMNLP Findings, 2024, pp. 567-578.
- R. Murmu, “Ol Chiki script formalization and standardization principles,” Santali Academy Journal, vol. 1, pp. 12-45, 1925.
- P. Koehn, Statistical Machine Translation, Cambridge University Press, 2009.
- J. Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers,” arXiv:1810.04805, 2018.
- A. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems, 2017.
- C. Raffel et al., “Exploring the limits of transfer learning with a unified text-to-text transformer,” arXiv:1910.10683, 2019.
- M. Lewis et al., “BART: Denoising sequence-to-sequence pre-training for natural language generation,” arXiv:1910.13461, 2019.
- J. Pk, “Santali language technology ecosystem: Production deployment and scaling strategies,” IEEE Kerala Section Conference, 2026.
To overcome the critical indigenous language extinction crisis affecting 7.6 million Santali speakers across eastern
India, Bangladesh, and Nepal, this research proposes a comprehensive AI-powered mobile application framework
systematically addressing systematic digital exclusion. Despite constitutional Eighth Schedule recognition (2003), Santali
confronts complete absence from commercial translation services (Google Translate, Microsoft Translator), negligible
Unicode Ol Chiki script implementation across 96% mobile platforms, and total lack of structured pedagogical digital
resources serving 42% illiterate demographic. The proposed Santali Translation and Tutor App implements real-time
bidirectional translation powered by Google Gemini 2.5 Flash model achieving state-of-the-art 32.6 BLEU score with 96.5%
native Ol Chiki script compliance through specialized three-tier prompt engineering cascade.
Keywords :
Santali Language Preservation, Ol Chiki Script Automation, Low-Resource Machine Translation, Flutter Crossplatform Development, CEFR-Aligned Mobile Pedagogy, Indigenous Language Revitalization, AI Prompt Engineering, Dialect Priority Systems.