Authors :
Vivina Vijesh; Dr. E. Gothai
Volume/Issue :
Volume 11 - 2026, Issue 7 - July
Google Scholar :
https://tinyurl.com/5n8mbt8t
Scribd :
https://tinyurl.com/3vwd3khh
DOI :
https://doi.org/10.38124/ijisrt/26jul962
Note : A published paper may take 4-5
working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and
ResearchGate.
Abstract :
Large Language Models (LLMs) have become indispensable across modern artificial intelligence applications,
yet their safety mechanisms continue to be evaluated almost exclusively on English inputs. In multilingual nations like
India, where millions communicate daily through Tanglish a fluid blend of Tamil and English this evaluation gap raises
fundamental concerns about whether these systems respond safely and consistently across diverse linguistic contexts. This
study addresses this critical oversight by introducing TanglishGuard, a novel benchmark designed to systematically assess
the robustness of LLM safety mechanisms against code-mixed inputs.
The benchmark evaluates four state-of-the-art models ChatGPT, Gemini, Claude, and DeepSeek using equivalent
harmful prompts expressed in English, Tamil, and Tanglish across nine distinct harm categories. Through rigorous
experimentation, our findings reveal that while all models demonstrate strong safety compliance on English and Tamil
inputs, Tanglish prompts reveal subtle but consistent vulnerabilities. These inconsistencies manifest as occasional failures
in detecting harmful intent within code-mixed language, highlighting significant gaps in multilingual AI safety
frameworks.
TanglishGuard provides a practical, reproducible framework for evaluating safety in mixed-language settings,
offering empirical evidence that current safety evaluations are insufficient for real-world multilingual usage. The
benchmark contributes to the development of more robust and equitable AI systems by ensuring safety mechanisms are
tested against authentic communication patterns rather than sanitised English-only datasets. This work underscores the
urgent imperative to move beyond English-centric safety evaluations. As AI systems become increasingly embedded in
diverse linguistic communities worldwide, ensuring their safety across the full spectrum of human language use is not
merely a technical challenge but a fundamental requirement for fairness, equity, and responsible AI deployment.
Keywords :
Tanglish, Code-Mixing, AI Safety, Large Language Models, Multilingual Benchmarks, Red-Teaming.
References :
- Wang, W., Tu, Z., Chen, C., Yuan, Y., Huang, J., Jiao, W., & Lyu, M. (2024). All Languages Matter: On the Multilingual Safety of LLMs. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 5865–5877). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.349
- Shen, L., Tan, W., Chen, S., Chen, Y., Zhang, J., Xu, H., Zheng, B., Koehn, P., & Khashabi, D. (2024). The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 2668–2680). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.156
- Yoo, H., Yang, Y., & Lee, H. (2025). Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 13392–13413). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.657
- Pattnayak, P., & Chowdhuri, S. (2026). IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia. arXiv preprint. https://arxiv.org/abs/2603.xxxxx
- Ning, Z., et al. (2025). LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models. arXiv preprint. https://arxiv.org/abs/2508.xxxxx
- Choi, D., et al. (2026). XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity. arXiv preprint. https://arxiv.org/abs/2605.05662
- OpenAI. (2023). GPT-4 Technical Report. arXiv preprint. https://doi.org/10.48550/arXiv.2303.08774
- Anthropic. (2022). Constitutional AI: Harmlessness from AI Feedback. Anthropic Research. https://www.anthropic.com/constitutional.pdf
- DeepSeek. (2025). DeepSeek-V3 Technical Report. arXiv preprint. https://arxiv.org/abs/2501.xxxxx
Large Language Models (LLMs) have become indispensable across modern artificial intelligence applications,
yet their safety mechanisms continue to be evaluated almost exclusively on English inputs. In multilingual nations like
India, where millions communicate daily through Tanglish a fluid blend of Tamil and English this evaluation gap raises
fundamental concerns about whether these systems respond safely and consistently across diverse linguistic contexts. This
study addresses this critical oversight by introducing TanglishGuard, a novel benchmark designed to systematically assess
the robustness of LLM safety mechanisms against code-mixed inputs.
The benchmark evaluates four state-of-the-art models ChatGPT, Gemini, Claude, and DeepSeek using equivalent
harmful prompts expressed in English, Tamil, and Tanglish across nine distinct harm categories. Through rigorous
experimentation, our findings reveal that while all models demonstrate strong safety compliance on English and Tamil
inputs, Tanglish prompts reveal subtle but consistent vulnerabilities. These inconsistencies manifest as occasional failures
in detecting harmful intent within code-mixed language, highlighting significant gaps in multilingual AI safety
frameworks.
TanglishGuard provides a practical, reproducible framework for evaluating safety in mixed-language settings,
offering empirical evidence that current safety evaluations are insufficient for real-world multilingual usage. The
benchmark contributes to the development of more robust and equitable AI systems by ensuring safety mechanisms are
tested against authentic communication patterns rather than sanitised English-only datasets. This work underscores the
urgent imperative to move beyond English-centric safety evaluations. As AI systems become increasingly embedded in
diverse linguistic communities worldwide, ensuring their safety across the full spectrum of human language use is not
merely a technical challenge but a fundamental requirement for fairness, equity, and responsible AI deployment.
Keywords :
Tanglish, Code-Mixing, AI Safety, Large Language Models, Multilingual Benchmarks, Red-Teaming.