Authors :
Tanu Chauhan; Abhishek Verma
Volume/Issue :
Volume 11 - 2026, Issue 8 - August
Google Scholar :
https://tinyurl.com/ycxdbsh8
DOI :
https://doi.org/10.38124/ijisrt/26aug324
Note : A published paper may take 4-5
working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and
ResearchGate.
Abstract :
Depression is a widespread and frequently under-diagnosed mental health condition, and delays in identification
are associated with poorer long-term outcomes. This paper presents a hybrid screening framework that combines a validated
questionnaire to the Patient Health Questionnaire-9 (PHQ-9), with linguistic feature extraction from short free-text journal
entries to generate a composite depression-risk score. Unlike approaches that rely on a single data modality, the proposed
system fuses structured clinical scoring with textual indicators associated with depressive symptomatology in prior
computational-linguistics research, such as elevated first-person pronoun usage, negation patterns, and lexical markers of
sadness, anhedonia, and worthlessness. A prototype web application implementing this framework was developed,
comprising a screening interface, a relational database schema for longitudinal storage, and an analytics dashboard for
aggregate monitoring. This paper describes the system architecture, the feature-fusion methodology, the underlying
database design, and a discussion of how the prototype's rule-based linguistic module can be replaced with a trained
supervised classifier (e.g., Support Vector Machine, Random Forest, or a fine-tuned transformer model) in future work. The
framework is intended as a low-cost, scalable screening aid to support — not replace — clinical [1] evaluation.
Keywords :
Depression Recognition; Machine Learning; Natural Language Processing; PHQ-9; Mental Health Screening; Text Classification; Linguistic Analysis.
References :
- M. De Choudhury, M. Gamon, S. Counts, and E. Horvitz, "Predicting depression via social media," in Proc. 7th Int. AAAI Conf. Weblogs and Social Media (ICWSM), 2013, pp. 128-137.
- S. Motade, A. Hassan, F. Mir, and K. Parikh, "Machine Learning-Based Approach for Depression Detection Using PHQ-9 and Twitter Dataset," in Proc. 3rd Int. Conf. Communication, Computing and Electronics Systems, Lecture Notes in Electrical Engineering, vol. 844, Springer, Singapore, 2022.
- Soumya Choudhary, Nikita Thomas, Janine Ellenberger, Grish Srinivasan , "A Machine Learning Approach for Detecting Digital Behavioral Patterns of Depression Using Nonintrusive Smartphone Data: Prospective Observational Study," JMIR Formative Research, 2022.
- Hadar Fisher, Nigel M. Jaffer , Kristina Pidcirny ,Anna O. Tierney Mia S. Vaidean , Poorvesh Dongre, and Christian A. Webb, "Language-based detection of depression with machine learning: systematic review and meta-analysis," npj Digital Medicine, 2026.
- Z. Shao, X. Wang, Z. Liu, C. Wang, and K. P. Subbalakshmi, "Systematic Evaluation of Machine-Generated Reasoning and PHQ-9 Labeling for Depression Detection Using Large Language Models," arXiv:2505.17119, 2025.
Depression is a widespread and frequently under-diagnosed mental health condition, and delays in identification
are associated with poorer long-term outcomes. This paper presents a hybrid screening framework that combines a validated
questionnaire to the Patient Health Questionnaire-9 (PHQ-9), with linguistic feature extraction from short free-text journal
entries to generate a composite depression-risk score. Unlike approaches that rely on a single data modality, the proposed
system fuses structured clinical scoring with textual indicators associated with depressive symptomatology in prior
computational-linguistics research, such as elevated first-person pronoun usage, negation patterns, and lexical markers of
sadness, anhedonia, and worthlessness. A prototype web application implementing this framework was developed,
comprising a screening interface, a relational database schema for longitudinal storage, and an analytics dashboard for
aggregate monitoring. This paper describes the system architecture, the feature-fusion methodology, the underlying
database design, and a discussion of how the prototype's rule-based linguistic module can be replaced with a trained
supervised classifier (e.g., Support Vector Machine, Random Forest, or a fine-tuned transformer model) in future work. The
framework is intended as a low-cost, scalable screening aid to support — not replace — clinical [1] evaluation.
Keywords :
Depression Recognition; Machine Learning; Natural Language Processing; PHQ-9; Mental Health Screening; Text Classification; Linguistic Analysis.