Authors :
Praise Elojo Attah; Lawrence Anebi Enyejo
Volume/Issue :
Volume 11 - 2026, Issue 8 - August
Google Scholar :
https://tinyurl.com/4pbamzh6
Scribd :
https://tinyurl.com/3zdedptj
DOI :
https://doi.org/10.38124/ijisrt/26aug331
Note : A published paper may take 4-5
working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and
ResearchGate.
Abstract :
Generative artificial intelligence systems increasingly operate as autonomous agents capable of retrieving external
information, accessing confidential resources, invoking application programming interfaces, and executing consequential
actions. These capabilities introduce substantial security risks because malicious instructions embedded in user prompts,
retrieved documents, webpages, emails, tool outputs, or persistent memory may alter an agent’s intended behaviour.
Conventional keyword filters, standalone prompt classifiers, and static access-control mechanisms provide limited
protection against semantically obfuscated attacks, multi-stage data exfiltration, manipulated tool arguments, and attacks
that remain within formally permitted privileges. This paper develops SecurePromptTrace, a novel runtime security
algorithm for detecting and controlling prompt injection, sensitive-data exfiltration, and tool misuse in generative artificial
intelligence systems.
SecurePromptTrace constructs a Dynamic Prompt Provenance Graph that represents trusted instructions, untrusted
content, model-generated plans, retrieved data, confidential variables, tool calls, tool arguments, and execution outcomes as
provenance-labelled nodes and causal edges. A relation-aware graph attention network analyses instruction dependencies
and identifies conflicts between the authenticated user objective and instructions originating from untrusted sources. The
graph model is integrated with a DeBERTa-v3 semantic injection classifier, confidential-data taint propagation, cross-layer
intent alignment, tool-capability compatibility analysis, and adaptive policy enforcement. A composite threat score combines
semantic injection probability, provenance conflict, sensitive-data flow, tool-privilege mismatch, execution-sequence
deviation, and predictive uncertainty. According to the calculated risk, the algorithm permits, sanitises, replans, isolates,
requests approval for, or blocks an operation.
The evaluation framework covers direct and indirect prompt injection, encoded and multilingual attacks, contextwindow attacks, memory poisoning, cross-tool exfiltration, parameter substitution, privilege chaining, and unauthorised
tool execution. SecurePromptTrace is compared with Aho–Corasick keyword filtering, regular-expression filtering, a
standalone DeBERTa-v3 classifier, PromptShield-style detection, role-based access control, attribute-based access control,
and combined static guardrails. Performance is assessed using macro-F1 score, precision, recall, AUROC, AUPRC, attack
success rate, data-exfiltration prevention rate, tool-misuse prevention rate, false-positive rate, task-utility retention,
computational latency, and memory overhead. Comparative bar charts, ROC and precision–recall curves, confusion
matrices, risk-score distributions, ablation graphs, and security–latency Pareto plots are used to demonstrate performance
differences. The central hypothesis is that provenance-aware semantic and behavioural tracing will produce significantly
lower attack-success and exfiltration rates than content-only or permission-only defenses while preserving legitimate task
completion. Statistical superiority will be established through confidence intervals, McNemar tests, bootstrap comparisons,
and effect-size analysis.
Keywords :
Prompt Injection; Data Exfiltration; Tool Misuse; Generative Artificial Intelligence; Static Access Controls.
References :
- Adewale, L. D. (2026). Digital evidence chains for PPAP assurance: AR-guided data capture, AI-verified documentation, and continuous audit automation for secure multi-tier supplier traceability in Industry 4.0 manufacturing. International Journal of Multidisciplinary Evolutionary Research, 7(1), 43–55. https://doi.org/10.54660/ijmer.2026.7.1.43-55
- Adewale, L. D. (2026). Smart factories, smarter evidence: Reinventing quality assurance for U.S. manufacturing competitiveness. International Journal of Multidisciplinary Futuristic Development, 7(1), 9–18. https://doi.org/10.54660/IJMFD.2026.7.1.09-18
- Agyekum, B., Seffah-Duodu, D., & Gloria, A. (2025). Federated industrial IoT threats detection. International Journal of Scientific Research and Modern Technology, 4(12), 226–244. https://doi.org/10.38124/ijsrmt.v4i12.1552
- Alotaibi, A., Mughus, R., & Ahmed, M. (2026). A red teaming framework for large language models: A case study on faithfulness evaluation. Software Quality Journal, 34, Article 42. https://doi.org/10.1007/s11219-026-09779-y
- Alzahrani, A. (2026). PromptGuard: A structured framework for injection-resilient language models. Scientific Reports, 16, Article 1277. https://doi.org/10.1038/s41598-025-31086-y
- Ayoola, V. B., Ugoaghalam, U. J., Idoko, I. P., Ijiga, O. M., and Olola, T. M. (2024). Effectiveness of social engineering awareness training in mitigating spear phishing risks in financial institutions from a cybersecurity perspective. Global Journal of Engineering and Technology Advances, 20(3), 94–117. https://doi.org/10.30574/gjeta.2024.20.3.0164
- Ayoola, V. B., Ugochukwu, U. N., Adeleke, I., Michael, C. I., Adewoye, M. B., and Adeyeye, Y. (2024). Generative AI-driven fraud detection in health care: Enhancing data loss prevention and cybersecurity analytics for real-time protection of patient records. International Journal of Scientific Research and Modern Technology, 3(11), 89–107. https://doi.org/10.38124/ijsrmt.v3i11.112
- Boiko, D. A., MacKnight, R., Kline, B., & Gomes, G. (2023). Autonomous chemical research with large language models. Nature, 624, 570–578. https://doi.org/10.1038/s41586-023-06792-0
- Briskinfosec, (2026). The Hidden Risk of Data Leakage in AI Code Assistants https://www.briskinfosec.com/blogs/blogsdetail/the-hidden-risk-of-data-leakage-in-ai-code-assistants
- Cai, Y. (2026). Prompt injection attacks on educational large language models for higher and vocational education. Scientific Reports, 16, Article 15594. https://doi.org/10.1038/s41598-026-46563-1
- Chen, K., Zhou, X., Lin, Y., Feng, S., Shen, L., & Wu, P. (2025). A survey on privacy risks and protection in large language models. Journal of King Saud University–Computer and Information Sciences, 37, Article 163. https://doi.org/10.1007/s44443-025-00177-1
- Clusmann, J., Ferber, D., Wiest, I. C., Schneider, C. V., Brinker, T. J., Foersch, S., Truhn, D., & Kather, J. N. (2025). Prompt injection attacks on vision language models in oncology. Nature Communications, 16, Article 1239. https://doi.org/10.1038/s41467-024-55631-x
- Debenedetti, E., Shumailov, I., Fan, T., Hayes, J., Carlini, N., Fabian, D., Kern, C., Shi, C., Terzis, A., & Tramèr, F. (2025). Defeating prompt injections by design. arXiv:2503.18813.
- Debenedetti, E., Zhang, J., Balunović, M., Beurer-Kellner, L., Fischer, M., & Tramèr, F. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. Advances in Neural Information Processing Systems, 37.
- Dzhaliuk, N., Sabodashko, D., Khoma, V., Khoma, Y., Kolchenko, V., and Podpora, M. (2026). Comparative evaluation of machine learning methods for protecting LLMs from prompt injection attacks. International Journal of Information Security, 25, Article 109. https://doi.org/10.1007/s10207-026-01264-8
- Gabla, E. S., Peter-Anyebe, A. C., & Ijiga, O. M. (2025). Assessing machine-learning-enabled anomaly detection models for real-time cyberattack mitigation in optical fiber communication systems. World Journal of Advanced Engineering Technology and Sciences, 17(2), 1–17. https://doi.org/10.30574/wjaets.2025.17.2.1454
- Hagendorff, T., Derner, E., & Oliver, N. (2026). Large reasoning models are autonomous jailbreak agents. Nature Communications, 17, Article 1435. https://doi.org/10.1038/s41467-026-69010-1
- He, P., Gao, J., & Chen, W. (2021). DeBERTaV3: Improving DeBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing. arXiv:2111.09543.
- Ibokette, A. I., Ogundare, T. O., Anyebe, A. P., Alao, F. O., Odeh, I. I., and Okafor, F. C. (2024). Mitigating maritime cybersecurity risks using AI-based intrusion detection systems and network automation during extreme environmental conditions. International Journal of Scientific Research and Modern Technology, 3(10), 65–91. https://doi.org/10.38124/ijsrmt.v3i10.73
- Idika, C. N., James, U. U., Ijiga, O. M., & Enyejo, L. A. (2023). Digital twin-enabled vulnerability assessment with Zero Trust policy enforcement in smart manufacturing cyber-physical systems. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 9(6), 475–499. https://doi.org/10.32628/CSEIT23906189
- Idoko, I. P., Ijiga, O. M., Enyejo, L. A., Akoh, O., Ebiega, G. I., Odeyemi, M. O., Olatunde, T. I., and Olajide, F. I. (2024). Integrating superhumans and synthetic humans into the Internet of Things and ubiquitous computing: Emerging AI applications and their relevance in the U.S. context. Global Journal of Engineering and Technology Advances, 19(1), 6–36. https://doi.org/10.30574/gjeta.2024.19.1.0055
- Igba, E., Olarinoye, H. S., Ezeh, N. V., Sehemba, D. B., Oluhaiyero, Y. S., & Okika, N. (2025). Synthetic data generation using generative AI to combat identity fraud and enhance global financial cybersecurity frameworks. International Journal of Scientific Research and Modern Technology, 4(2), 1–19. https://doi.org/10.5281/zenodo.14928919
- Ijiga, O. M., Idoko, I. P., Ebiega, G. I., Olajide, F. I., Olatunde, T. I., and Ukaegbu, C. (2024). Harnessing adversarial machine learning for advanced threat detection: AI-driven strategies in cybersecurity risk assessment and fraud prevention. Open Access Research Journal of Science and Technology, 11(1), 1–24. https://doi.org/10.53022/oarjst.2024.11.1.0060
- Ijiga, O. M., Okika, N., Balogun, S. A., Enyejo, L. A., and Agbo, O. J. (2025). A comprehensive review of federated learning architectures for insider threat detection in distributed SQL-based enterprise environments. International Journal of Innovative Science and Research Technology, 10(7), 536–550. https://doi.org/10.38124/ijisrt/25jul392
- Jalan, P., Abishethvarman, V., Chandna, B., & Naseem, U. (2026). Survey on LLM safety: Attacks, defenses, alignment, metrics, and guardrails. Machine Learning, 115, Article 130. https://doi.org/10.1007/s10994-026-07060-8
- James, U. U. (2026). Multiscale Signal Processing Framework for Zero Trust Security in Next- Generation Wireless Networks International Journal of Wireless and Mobile Networks DOI/Link: https://aircconline.com/ijwmn/V18N3/18326ijwmn01.pdf
- James, U. U., and Akujuobi, C. M. (2026). Multiscale signal processing framework for Zero Trust security in next-generation wireless networks. International Journal of Wireless and Mobile Networks, 18(3), 1–21. https://doi.org/10.5121/ijwmn.2026.18301
- James, U. U., Olarinoye, H. S., Uchenna, I. R., Idika, C. N., Ngene, O. J., Ijiga, O. M., and Itemuagbor, K. (2025). Combating deepfake threats using X-FACTS explainable CNN framework for enhanced detection and cybersecurity resilience. Advances in Artificial Intelligence and Robotics Research, 1, 41–64. https://doi.org/10.4236/airr.2025.11004
- James, U. U., Olarinoye, H. S., Uchenna, I. R., Idika, C. N., Ngene, O. J., Ijiga, O. M., & Itemuagbor, K. (2025). Combating deepfake threats using X-FACTS explainable CNN framework for enhanced detection and cybersecurity resilience. Advances in Artificial Intelligence and Robotics Research, 1, 41–64. https://doi.org/10.4236/airr.2025.11004
- James, U. U., Salami, E. O., & Enyejo, L. A. (2025). Real-time policy orchestration for cybersecurity risk management in GRC-aligned financial technology infrastructures. International Journal of Innovative Science and Research Technology, 10(8). https://doi.org/10.38124/ijisrt/25aug1021
- James, U. U., Salami, E. O., & Enyejo, L. A. (2025). Real-time policy orchestration for cybersecurity risk management in GRC-aligned financial technology infrastructures. International Journal of Innovative Science and Research Technology, 10(8). https://doi.org/10.38124/ijisrt/25aug1021
- James, U. U., Salami, E. O., & Enyejo, L. A. (2025). Real-time policy orchestration for cybersecurity risk management in GRC-aligned financial technology infrastructures. International Journal of Innovative Science and Research Technology, 10(8). https://doi.org/10.38124/ijisrt/25aug1021
- Ko, R., Chung, J., Zheng, S., Xiao, C., Kim, T.-W., Onizuka, M., & Shin, W.-Y. (2026). Seven security challenges in cross-domain multi-agent LLM systems. npj Artificial Intelligence. https://doi.org/10.1038/s44387-026-00128-9
- Kpogli, S. A., Onwuzurike, M. A., & Ayoola, V. B. (2024). Optimizing Educational Return on Investment Through AI-Driven Curriculum Adaptation and Performance-Based Resource Allocation Models. International Journal of Scientific Research and Modern Technology, 3(9), 141–156. https://doi.org/10.38124/ijsrmt.v3i9.1352
- Kpogli, S. A., Onwuzurike, M. A., & Enyejo, J. O. (2024). Integrating artificial intelligence and learning sciences to reduce cognitive load and achievement gaps in data-driven K–12 instructional systems. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 10(6), 2569–2589. https://doi.org/10.32628/CSEIT25113575
- Kwarteng, R. A., Idoko, I. P., Ijiga, O. M., & Enyejo, L. A. (2020). Integrating cybersecurity awareness and access control into organizational IT operations for risk reduction. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 6(1), 243–261. https://doi.org/10.32628/CSEIT23906128
- Milani, A., Franzoni, V., & Florindi, E. (2026). Indirect prompt injection in large language models. Neural Computing and Applications, 38, Article 530. https://doi.org/10.1007/s00521-026-12266-x
- Okika, N., Okoh, O. F., & Etuk, E. E. (2025). Mitigating insider threats and social engineering tactics in advanced persistent threat operations through behavioral analytics and cybersecurity training. International Journal of Advance Research Publication and Reviews, 2(3), 11–27.
- Ononiwu, M., Azonuche, T. I., Imoh, P. O., & Enyejo, J. O. (2024). Evaluating blockchain content monetization platforms for autism-focused streaming with cybersecurity and scalable microservice architectures. Iconic Research and Engineering Journals, 8(1), 722–740.
- Onwuzurike, M. A. & Enyejo, J. O. (2026). A Business Intelligence Framework for AI Powered Educational Platforms Linking Learning Analytics to Strategic Decision Making in K-12 Schools International Journal of Recent Research in Commerce Economics and Management (IJRRCEM) Vol. 13, Issue 2, pp: (21-42), DOI: https://doi.org/ 10.5281/zenodo.19510038
- Onwuzurike, M. A. (2023). Human-Centered Design of Intelligent Tutoring Systems Integrating Behavioral Analytics and Inclusive Pedagogical Principles for Early Learners International Journal of Scientific Research in Science, Engineering and Technology Volume 10, Issue 3, Page Number 720-738, doi : https://doi.org/10.32628/IJSRSET2310330
- Onwuzurike, M. A., & Kpogli, S. A. (2022). Data-informed strategic management of EdTech startups leveraging artificial intelligence for sustainable K–12 learning innovation. International Journal of Scientific Research and Modern Technology, 1(12), 187–200. https://doi.org/10.38124/ijsrmt.v1i12.1117
- Onwuzurike, M. A., & Kpogli, S. A. (2025). Predictive modeling of student engagement and behavioral outcomes using machine learning techniques in technology-enhanced classrooms. International Journal of Scientific Research in Humanities and Social Sciences, 2(6), 58–79. https://doi.org/10.32628/IJSRHSS2525135
- Onwuzurike, M. A., & Raphael, F. O. (2025). Ethical governance models for artificial intelligence deployment in K–12 education: Balancing algorithmic personalization, accountability and child protection policy. International Journal of Scientific Research and Modern Technology, 4(8), 193–208. https://doi.org/10.38124/ijsrmt.v4i8.1271
- Onwuzurike, M. A., Enyejo, J. O., & Peter-Anyebe, A. C. (2026). Design and evaluation of real-time adaptive learning algorithms for personalized K–12 curriculum optimization using student performance analytics. World Journal of Advance Multidisciplinary Research, 3(3), 21–36. https://doi.org/10.5281/zenodo.19131296
- Onwuzurike, M. A., Igba, E. (2023). Applying explainable machine learning models to educational data for transparent decision support in curriculum design and student assessment. Journal of Frontiers in Multidisciplinary Research. 2023;4(1):585–599. doi:10.54660/.JFMR.2023.4.1.585-599
- Shilov, I., Meeus, M., & de Montjoye, Y.-A. (2026). The mosaic memory of large language models. Nature Communications, 17, Article 2142. https://doi.org/10.1038/s41467-026-68603-0
- Srinivasan, J., Regi, S. A., Anbarasan, A. K., Suresh, A., Vetriselvi, T., & Venu, S. (2026). Detection and analysis of prompt injection in Indian multilingual large language models. Scientific Reports, 16, Article 16208. https://doi.org/10.1038/s41598-026-43883-0
- Ussher-Eke, D., Onoja, D. A., Ijiga, O. M., & Enyejo, L. A. (2025). Strengthening human resource compliance and ethical oversight through cybersecurity awareness and policy enforcement. International Journal of Management and Commerce Innovations, 13(1), 376–393. https://doi.org/10.5281/zenodo.16539315
- Wang, W., Liu, J., Gu, B., Wang, C., Qu, Y. N., & Feng, D. (2026). Securing LLM-in-the-loop software for empirical study of risks, mitigations, and utility trade-offs in a safety-critical case. Empirical Software Engineering, 31(4), Article 87. https://doi.org/10.1007/s10664-026-10820-8
- Webb, T., Mondal, S. S., & Momennejad, I. (2025). A brain-inspired agentic architecture to improve planning with LLMs. Nature Communications, 16, Article 8633. https://doi.org/10.1038/s41467-025-63804-5
- Yang, Y., Jin, Q., Huang, F., and Lu, Z. (2025). Adversarial prompt and fine-tuning attacks threaten medical large language models. Nature Communications, 16, Article 9011. https://doi.org/10.1038/s41467-025-64062-1
Generative artificial intelligence systems increasingly operate as autonomous agents capable of retrieving external
information, accessing confidential resources, invoking application programming interfaces, and executing consequential
actions. These capabilities introduce substantial security risks because malicious instructions embedded in user prompts,
retrieved documents, webpages, emails, tool outputs, or persistent memory may alter an agent’s intended behaviour.
Conventional keyword filters, standalone prompt classifiers, and static access-control mechanisms provide limited
protection against semantically obfuscated attacks, multi-stage data exfiltration, manipulated tool arguments, and attacks
that remain within formally permitted privileges. This paper develops SecurePromptTrace, a novel runtime security
algorithm for detecting and controlling prompt injection, sensitive-data exfiltration, and tool misuse in generative artificial
intelligence systems.
SecurePromptTrace constructs a Dynamic Prompt Provenance Graph that represents trusted instructions, untrusted
content, model-generated plans, retrieved data, confidential variables, tool calls, tool arguments, and execution outcomes as
provenance-labelled nodes and causal edges. A relation-aware graph attention network analyses instruction dependencies
and identifies conflicts between the authenticated user objective and instructions originating from untrusted sources. The
graph model is integrated with a DeBERTa-v3 semantic injection classifier, confidential-data taint propagation, cross-layer
intent alignment, tool-capability compatibility analysis, and adaptive policy enforcement. A composite threat score combines
semantic injection probability, provenance conflict, sensitive-data flow, tool-privilege mismatch, execution-sequence
deviation, and predictive uncertainty. According to the calculated risk, the algorithm permits, sanitises, replans, isolates,
requests approval for, or blocks an operation.
The evaluation framework covers direct and indirect prompt injection, encoded and multilingual attacks, contextwindow attacks, memory poisoning, cross-tool exfiltration, parameter substitution, privilege chaining, and unauthorised
tool execution. SecurePromptTrace is compared with Aho–Corasick keyword filtering, regular-expression filtering, a
standalone DeBERTa-v3 classifier, PromptShield-style detection, role-based access control, attribute-based access control,
and combined static guardrails. Performance is assessed using macro-F1 score, precision, recall, AUROC, AUPRC, attack
success rate, data-exfiltration prevention rate, tool-misuse prevention rate, false-positive rate, task-utility retention,
computational latency, and memory overhead. Comparative bar charts, ROC and precision–recall curves, confusion
matrices, risk-score distributions, ablation graphs, and security–latency Pareto plots are used to demonstrate performance
differences. The central hypothesis is that provenance-aware semantic and behavioural tracing will produce significantly
lower attack-success and exfiltration rates than content-only or permission-only defenses while preserving legitimate task
completion. Statistical superiority will be established through confidence intervals, McNemar tests, bootstrap comparisons,
and effect-size analysis.
Keywords :
Prompt Injection; Data Exfiltration; Tool Misuse; Generative Artificial Intelligence; Static Access Controls.