Authors :
Ahmed F. Mohamed
Volume/Issue :
Volume 11 - 2026, Issue 7 - July
Google Scholar :
https://tinyurl.com/mr28aa2w
Scribd :
https://tinyurl.com/fyjmztms
DOI :
https://doi.org/10.38124/ijisrt/26jul649
Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.
Abstract :
Production data migrations run with write credentials, often while the application they serve continues to handle
traffic, and their worst failure modes concern how they change data rather than whether the code runs. Language model
reviewers are increasingly asked to gate such scripts, with little evidence about their reliability in this setting. This paper
presents MigBench, a benchmark of 300 MongoDB migration scripts in which 100 are correct and 200 each contain
exactly one defect from eight operationally defined categories. The dataset is generated deterministically from a single
seed, and every label is certified by execution: each script runs against a disposable MongoDB replica set under five
behavioral probes covering expected state and scope, repeated execution, a counter race against simulated live traffic,
crash injection with an invariant across collections, and crash injection followed by resume. All 300 labels were confirmed
by behavior before any reviewer ran.
Keywords :
Data Migrations; MongoDB; Large Language Models; Code Review; Automated Program Repair; Benchmarks; Fault Injection; Multi-Agent Systems.
References :
- Z. Li et al., “Automating code review activities by large-scale pre-training,” in Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE), 2022. doi: 10.1145/3540250.3549081.
- R. Tufano, S. Masiero, A. Mastropaolo, L. Pascarella, D. Poshyvanyk, and G. Bavota, “Using pre-trained models to boost code review automation,” in Proceedings of the 44th International Conference on Software Engineering (ICSE), 2022. arXiv:2201.06850.
- C. E. Jimenez et al., “SWE-bench: Can language models resolve real-world GitHub issues?” in Proceedings of the 12th International Conference on Learning Representations (ICLR), 2024. arXiv:2310.06770.
- J. Yang et al., “SWE-agent: Agent-computer interfaces enable automated software engineering,” in Advances in Neural Information Processing Systems 37 (NeurIPS), 2024. arXiv:2405.15793.
- R. Just, D. Jalali, and M. D. Ernst, “Defects4J: A database of existing faults to enable controlled testing studies for Java programs,” in Proceedings of the International Symposium on Software Testing and Analysis (ISSTA), 2014. doi: 10.1145/2610384.2628055.
- M. Chen et al., “Evaluating large language models trained on code,” arXiv preprint arXiv:2107.03374, 2021.
- J. Liu, C. S. Xia, Y. Wang, and L. Zhang, “Is your code generated by ChatGPT really correct? Rigorous evaluation of large language models for code generation,” in Advances in Neural Information Processing Systems 36 (NeurIPS), 2023. arXiv:2305.01210.
- S. Hong et al., “MetaGPT: Meta programming for a multi-agent collaborative framework,” in Proceedings of the 12th International Conference on Learning Representations (ICLR), 2024. arXiv:2308.00352.
- C. Qian et al., “ChatDev: Communicative agents for software development,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024, pp. 15174–15186. doi: 10.18653/v1/2024.acl-long.810.
- Q. Wu et al., “AutoGen: Enabling next-gen LLM applications via multi-agent conversations,” in Proceedings of the First Conference on Language Modeling (COLM), 2024. arXiv:2308.08155.
- J. Li, Q. Zhang, Y. Yu, Q. Fu, and D. Ye, “More agents is all you need,” Transactions on Machine Learning Research, 2024. arXiv:2402.05120.
- M. Cemri et al., “Why do multi-agent LLM systems fail?” arXiv preprint arXiv:2503.13657, 2025.
- L. Zheng et al., “Judging LLM-as-a-judge with MT-Bench and Chatbot Arena,” in Advances in Neural Information Processing Systems 36 (NeurIPS), Datasets and Benchmarks Track, 2023. arXiv:2306.05685.
- P. Wang et al., “Large language models are not fair evaluators,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024, pp. 9440–9450. doi: 10.18653/v1/2024.acl-long.511.
- C. A. Curino, H. J. Moon, and C. Zaniolo, “Graceful database schema evolution: The PRISM workbench,” Proceedings of the VLDB Endowment, vol. 1, no. 1, pp. 761–772, 2008. doi: 10.14778/1453856.1453939.
- D. Qiu, B. Li, and Z. Su, “An empirical analysis of the co-evolution of schema and code in database applications,” in Proceedings of the 9th Joint Meeting on Foundations of Software Engineering (ESEC/FSE), 2013, pp. 125–135. doi: 10.1145/2491411.2491431.
- S. Scherzinger, M. Klettke, and U. Störl, “Managing schema evolution in NoSQL data stores,” in Proceedings of the 14th International Symposium on Database Programming Languages (DBPL), 2013. arXiv:1308.0514.
- U. Störl, M. Klettke, and S. Scherzinger, “NoSQL schema evolution and data migration: State-of-the-art and opportunities,” in Proceedings of the 23rd International Conference on Extending Database Technology (EDBT), 2020, pp. 655–658. doi: 10.5441/002/edbt.2020.87.
- S. W. Ambler and P. J. Sadalage, Refactoring Databases: Evolutionary Database Design. Boston, MA: Addison-Wesley, 2006.
- T. Yu et al., “Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2018, pp. 3911–3921. doi: 10.18653/v1/D18-1425.
- J. Li et al., “Can LLM already serve as a database interface? A big bench for large-scale database grounded text-to-SQLs,” in Advances in Neural Information Processing Systems 36 (NeurIPS), Datasets and Benchmarks Track, 2023. arXiv:2305.03111.
- X. Peng, Y. Zhang, J. Yang, and M. Stevenson, “On the vulnerabilities of text-to-SQL models,” in Proceedings of the 34th IEEE International Symposium on Software Reliability Engineering (ISSRE), 2023. doi: 10.1109/ISSRE59848.2023.00047.
- P. Bailis, A. Fekete, M. J. Franklin, A. Ghodsi, J. M. Hellerstein, and I. Stoica, “Feral concurrency control: An empirical investigation of modern application integrity,” in Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data, 2015, pp. 1327–1342. doi: 10.1145/2723372.2737784.
- T. Warszawski and P. Bailis, “ACIDRain: Concurrency-related attacks on database-backed web applications,” in Proceedings of the 2017 ACM SIGMOD International Conference on Management of Data, 2017. doi: 10.1145/3035918.3064037.
- R. A. DeMillo, R. J. Lipton, and F. G. Sayward, “Hints on test data selection: Help for the practicing programmer,” Computer, vol. 11, no. 4, pp. 34–41, 1978. doi: 10.1109/C-M.1978.218136.
- Y. Jia and M. Harman, “An analysis and survey of the development of mutation testing,” IEEE Transactions on Software Engineering, vol. 37, no. 5, pp. 649–678, 2011. doi: 10.1109/TSE.2010.62.
- G. Ramalingam and K. Vaswani, “Fault tolerance via idempotence,” in Proceedings of the 40th ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages (POPL), 2013, pp. 249–262. doi: 10.1145/2429069.2429100.
- P. Alvaro, J. Rosen, and J. M. Hellerstein, “Lineage-driven fault injection,” in Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data, 2015. doi: 10.1145/2723372.2723711.
- Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,” Psychometrika, vol. 12, no. 2, pp. 153–157, 1947. doi: 10.1007/BF02295996.
- T. G. Dietterich, “Approximate statistical tests for comparing supervised classification learning algorithms,” Neural Computation, vol. 10, no. 7, pp. 1895–1923, 1998. doi: 10.1162/089976698300017197.
- A. Eghbali and M. Pradel, “De-Hallucinator: Mitigating LLM hallucinations in code generation tasks via iterative grounding,” arXiv preprint arXiv:2401.01701, 2024.
32. J. Spracklen, R. Wijewickrama, A. H. M. N. Sakib, A. Maiti, B. Viswanath, and M. Jadliwala, “We have a package for you! A comprehensive analysis of package hallucinations by code generating LLMs,” in Proceedings of the 34th USENIX Security Symposium, 2025. arXiv:2406.10279.
Production data migrations run with write credentials, often while the application they serve continues to handle
traffic, and their worst failure modes concern how they change data rather than whether the code runs. Language model
reviewers are increasingly asked to gate such scripts, with little evidence about their reliability in this setting. This paper
presents MigBench, a benchmark of 300 MongoDB migration scripts in which 100 are correct and 200 each contain
exactly one defect from eight operationally defined categories. The dataset is generated deterministically from a single
seed, and every label is certified by execution: each script runs against a disposable MongoDB replica set under five
behavioral probes covering expected state and scope, repeated execution, a counter race against simulated live traffic,
crash injection with an invariant across collections, and crash injection followed by resume. All 300 labels were confirmed
by behavior before any reviewer ran.
Keywords :
Data Migrations; MongoDB; Large Language Models; Code Review; Automated Program Repair; Benchmarks; Fault Injection; Multi-Agent Systems.