AI-Driven Benchmark and Evaluation Framework for Incident Diagnosis in Cloud-Scale Systems | IJCT Volume 13 – Issue 4 | IJCT-V13I4P17

IJCT
International Journal of Computer Techniques
ISSN 2394-2231 · Peer-Reviewed · Open Access
📚 Volume 13, Issue 4
📅 August 8, 2026
📄 Pages 165–176
🔖 ID: IJCT-V13I4P17

AI-Driven Benchmark and Evaluation Framework for Incident Diagnosis in Cloud-Scale Systems

Author(s)

Venkata Praveen Annam

Abstract

The exponential growth of microservice architectures and cloud-native deployments has fundamentally transformed the landscape of incident management. As organizations operate platforms encompassing thousands of interdependent services, the challenge of rapid, accurate incident diagnosis has become a critical engineering problem. This paper presents the AI-Driven Diagnostic Benchmark Framework (AI-DBF), a comprehensive evaluation system designed to rigorously assess AI-based incident diagnosis tools across cloud-scale environments. AI-DBF introduces a curated corpus of 4,812 labeled production incidents spanning six canonical failure categories, a reproducible fault-injection harness targeting Kubernetes-based deployments, and a multi-dimensional scoring protocol evaluating detection accuracy, diagnosis latency, root-cause attribution, and graceful degradation under telemetry noise. Experimental evaluation across four representative systems demonstrates that AI-DBF exposes meaningful performance differentials invisible to existing ad hoc assessments. Our framework achieves a mean-time-to-detect (MTTD) of 6.2 minutes, an F1 score of 0.941 for incident classification, and sustains sub-10-second P95 diagnosis latency at 5,000-service scale. AI-DBF is designed to be vendor-neutral, reproducible, and continuously extensible, providing the research community with a shared foundation for fair, evidence-based comparison of emerging diagnostic AI systems.

Keywords

cloud incident management, AI benchmark framework, fault injection, root cause analysis, microservices, observability, SRE automation, anomaly detection

Conclusion

This paper has presented AI-DBF, a comprehensive benchmark and evaluation framework for AI-driven incident diagnosis in cloud-scale systems. Through a curated corpus of 4,812 labeled production incidents, a reproducible fault-injection harness, and a seven-dimensional evaluation protocol, AI-DBF provides the research community with the shared evaluation infrastructure needed to conduct fair, rigorous, and reproducible comparisons of incident diagnosis systems. Empirical evaluation of four representative systems demonstrates that AI-DBF exposes meaningful performance differentials across all evaluation dimensions, with the performance gaps between systems being most pronounced on noise robustness and scalability dimensions that existing ad hoc evaluations typically ignore. The AI-DBF integrated diagnosis system achieves an F1 of 0.941, MTTD of 6.2 minutes, and P95 diagnosis latency of 8.7 seconds, establishing concrete performance targets for future work. We release the AI-DBF incident corpus, fault-injection harness, evaluation code, and baseline system implementations as open-source artifacts. We hope this work contributes not only to the methodological rigor of cloud incident diagnosis research, but ultimately to the practical goal of reducing the human cost of operating complex, large-scale systems.

References

[1] Ahmed, T., Bhattacharya, S., & Goel, A. (2019). Robust anomaly detection in cloud microservices using multivariate time-series clustering. In Proceedings of the 14th ACM International Conference on Distributed and Event-based Systems (DEBS), pp. 112–123. ACM.
[2] Chen, P., Qi, Y., Zheng, P., & Hou, D. (2020). CausInfer: Automatic and distributed performance diagnosis with hierarchical causality graph in large distributed systems. In IEEE INFOCOM 2020, pp. 1704–1713. IEEE.
[3] Cigna, L., Klues, T., & Muñoz-Gea, J. (2021). Lessons from five years of AI-assisted incident response at Netflix. Netflix Engineering Blog. Retrieved from https://netflixtechblog.com.
[4] Du, M., Li, F., Zheng, G., & Srikumar, V. (2017). DeepLog: Anomaly detection and diagnosis from system logs through deep learning. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security (CCS), pp. 1285–1298.
[5] Farshchi, M., Schneider, J.-G., Weber, I., & Grundy, J. (2018). Experience report: Anomaly detection of cloud application operations using log and cloud metric correlation analysis. In Proceedings of the 28th International Symposium on Software Reliability Engineering (ISSRE), pp. 24–35. IEEE.
[6] Beyer, B., Jones, C., Petoff, J., & Murphy, N. R. (Eds.). (2016). Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media.
[7] Li, Z., Chen, P., Hou, D., Sun, Y., & Liao, Q. (2021). Practical root cause localization for microservice systems via trace analysis. In Proceedings of the 2021 IEEE/ACM 29th International Symposium on Quality of Service (IWQOS), pp. 1–10.
[8] Luo, C., Du, J., Ye, C., Li, H., & Shan, Z. (2022). AIOps challenges and opportunities in site reliability engineering: Case studies at Alibaba Cloud. IEEE Transactions on Services Computing, 15(4), 2276–2291.
[9] Meng, Y., Zhang, S., Sun, Y., Zhang, R., Hu, Z., Zhang, Y., & Pei, D. (2020). Localizing failure root causes in a microservice through causality inference. In 2020 IEEE/ACM 28th International Symposium on Quality of Service (IWQoS), pp. 1–10.
[10] Nedelkoski, S., Bogatinovski, J., Acker, A., Cardoso, J., & Kao, O. (2021). Self-supervised log parsing. In Machine Learning and Knowledge Discovery in Databases (ECML PKDD), pp. 122–138. Springer.
[11] He, P., Zhu, J., Zheng, Z., & Lyu, M. R. (2016). Drain: An online log parsing approach with fixed depth tree. In 2016 IEEE International Conference on Web Services (ICWS), pp. 33–40.
[12] Zhao, N., Chen, P., Zheng, Z., & Sui, K. (2020). Automatically and adaptively identifying severe alerts for online service systems. In Proceedings of IEEE INFOCOM 2020, pp. 2445–2454.
[13] Botezatu, M. M., Giurgiu, I., Bogojeska, J., & Wiesmann, D. (2016). Predicting disk replacement towards reliable data centers. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 39–48.
[14] Zhang, W., Xu, Y., Lin, X., Zhang, Z., & Wang, J. (2024). RCACopilot: On-call empowerment for cloud incident root cause analysis with large language models. In Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering (FSE), pp. 1802–1814.
[15] Dang, Y., Lin, Q., & Huang, P. (2019). AIOps: Real-world challenges and research innovations. In Proceedings of the 41st International Conference on Software Engineering: Companion (ICSE-Companion), pp. 4–5. IEEE.

📋 How to Cite This Paper

Venkata Praveen Annam (2026). AI-Driven Benchmark and Evaluation Framework for Incident Diagnosis in Cloud-Scale Systems. International Journal of Computer Techniques, 13(4), 165–176. ISSN: 2394-2231. DOI: https://doi.org/10.5281/zenodo.21853784
© 2026 International Journal of Computer Techniques (IJCT). All rights reserved. · ijctjournal.org
Submit Your Paper