Authors :
Subodh Kant; Dr. Gaurav Harit
Volume/Issue :
Volume 11 - 2026, Issue 7 - July
Google Scholar :
https://tinyurl.com/52wm8eau
Scribd :
https://tinyurl.com/2kzrcpdk
DOI :
https://doi.org/10.38124/ijisrt/26jul106
Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.
Abstract :
This survey presents a comprehensive synthesis of contemporary research in handwritten text recognition
(HTR), encompassing bibliometric analysis, systematic methodological reviews, and novel architectural innovations
spanning line-level, document-level, and multi-lingual recognition systems. Drawing from 25 original research works,
survey papers, and empirical studies, this work categories existing literature into ten thematic clusters: survey and review
studies, meta-learning and adaptation methods, self-supervised learning approaches, vision-language models, transformerbased architectures, word and keyword spotting methods, document-level recognition systems, low-resource and Indic
script recognition, foundational deep learning architectures, and out-of-distribution generalization studies. The survey
reveals that while significant progress has been achieved through deep learning and transfer learning techniques, critical
challenges persist in handling domain shifts, low-resource languages, and complex document layouts. Furthermore,
emerging paradigms including self-supervised vision transformers, meta-learning frameworks, and foundation models
demonstrate substantial promise for enabling more adaptive, efficient, and generalisable HTR systems. This paper
synthesises these developments, identifies cross-cutting methodological themes, analyses comparative performance trends,
and delineates key research gaps that warrant future investigation.
Keywords :
Handwritten Text Recognition, Deep Learning, Meta-Learning, Vision Transformers, Self-Supervised Learning, Domain Adaptation, Foundation Models, Indic Scripts.
References :
- V. Agrawal, J. Jagtap, and M. V. V. Prasad Kantipudi, “Exploration of advancements in handwritten document recognition techniques,” Intelligent Systems with Applications, vol. 22, Art. no. 200346, 2024.
- C. Garrido-Muñoz, A. Rios-Vila, and J. Calvo-Zaragoza, “Handwritten text recognition: A survey,” arXiv preprint arXiv:2502.08417, 2025.
- A. K. Bhunia, S. Khan, A. Fischer, F. S. Khan, and L. Shao, “MetaHTR: Meta-learning for writer-adaptive handwritten text recognition,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 1581–1590.
- W. Gu, L. Gu, Z. Wang, C. Y. Suen, and Y. Wang, “DocTTT: Test-Time Training for Handwritten Document Recognition Using Meta-Auxiliary Learning,” arXiv preprint arXiv:2501.12898, 2025.
- W. Gu, L. Gu, C. Y. Suen, and Y. Wang, “MetaWriter: Personalized handwritten text recognition using meta-learned prompt tuning,” arXiv preprint arXiv:2505.20513, 2025.
- S. K. Jemni, M. E. Souibgui, Y. Kessentini, and A. Fornés, “ST-KeyS: Self-supervised transformer for keyword spotting in historical documents,” in Proc. Digital Humanities Conference, 2023.
- F. Wolf and G. A. Fink, “Self-training of handwritten word recognition for synthetic-to-real adaptation,” in Proc. International Conference on Document Analysis and Recognition (ICDAR), 2019, pp. 1–8.
- M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, et al., “Emerging properties in self-supervised vision transformers,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 9650–9660.
- M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, et al., “DINOv2: Learning robust visual features without supervision,” arXiv preprint arXiv:2304.07193, 2023.
- Y. H. M. Chung and D. Choi, “Finetuning vision-language models as OCR systems for low-resource languages: A case study of Manchu,” arXiv preprint arXiv:2507.06761, 2025.
- A. Fadeeva, V. Coriou, D. Antognini, C. Musat, and A. Maksai, “InkFM: A foundational model for full-page online handwritten note understanding,” arXiv preprint arXiv:2503.23081, 2025.
- A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, et al., “Learning transferable visual models from natural language supervision,” in Proc. International Conference on Machine Learning (ICML), 2021, pp. 8748–8763.
- Y. Li, Y. Zhu, and Z. Lian, “HTR-VT: Handwritten text recognition with vision transformer,” arXiv preprint arXiv:2409.08573, 2024.
- L. Kumari, P. Mazumdar, and P. Bhattacharyya, “GatedLexiconNet: End-to-end handwritten paragraph recognition,” arXiv preprint arXiv:2404.14062, 2024.
- M. Hamdan, L. S. Saoud, and N. Houmani, “HTR-JAND: Joint attention network and knowledge distillation for handwritten text recognition,” arXiv preprint arXiv:2412.18524, 2024.
- D. Coquenet, “Meta-DAN: Towards an efficient prediction strategy for page-level handwritten text recognition,” arXiv preprint arXiv:2504.03349, 2025.
- B. V. Kasuba, M. K. Chinnakotla, and C. V. Jawahar, “PLATTER: Page-level recognition for Indic languages,” arXiv preprint arXiv:2502.06172, 2025.
- L. Sharma, A. Bhatt, and N. Garg, “Advancements in handwritten Devanagari character recognition using deep learning,” Discover Applied Sciences, vol. 6, 2024.
- K. V. C. Madhav, A. Veeramaneni, P. A. Sriya, K. Sumanth, J. C. Jyotsna, T. Singh, and P. Duraisamy, “Handwritten Devanagari numeral recognition using deep learning,” in Proc. 2024 3rd International Conference for Innovation in Technology (INOCON), Karnataka, India, Mar. 2024, pp. 1–6, doi: 10.1109/INOCON60754.2024.10511947.
- C. Garrido-Muñoz and J. Calvo-Zaragoza, “On the generalization of handwritten text recognition models,” arXiv preprint arXiv:2411.17332, 2024.
This survey presents a comprehensive synthesis of contemporary research in handwritten text recognition
(HTR), encompassing bibliometric analysis, systematic methodological reviews, and novel architectural innovations
spanning line-level, document-level, and multi-lingual recognition systems. Drawing from 25 original research works,
survey papers, and empirical studies, this work categories existing literature into ten thematic clusters: survey and review
studies, meta-learning and adaptation methods, self-supervised learning approaches, vision-language models, transformerbased architectures, word and keyword spotting methods, document-level recognition systems, low-resource and Indic
script recognition, foundational deep learning architectures, and out-of-distribution generalization studies. The survey
reveals that while significant progress has been achieved through deep learning and transfer learning techniques, critical
challenges persist in handling domain shifts, low-resource languages, and complex document layouts. Furthermore,
emerging paradigms including self-supervised vision transformers, meta-learning frameworks, and foundation models
demonstrate substantial promise for enabling more adaptive, efficient, and generalisable HTR systems. This paper
synthesises these developments, identifies cross-cutting methodological themes, analyses comparative performance trends,
and delineates key research gaps that warrant future investigation.
Keywords :
Handwritten Text Recognition, Deep Learning, Meta-Learning, Vision Transformers, Self-Supervised Learning, Domain Adaptation, Foundation Models, Indic Scripts.