
Incident Knowledge Graphs for Site Reliability Engineering: Connecting Alerts, Runbooks, Services, Deployments, and Postmortems | IJCT Volume 13 – Issue 5 | IJCT-V13I5P48
IJCT
International Journal of Computer Techniques
ISSN 2394-2231 · Peer-Reviewed · Open Access
📚 Volume 13, Issue 5
📅 September 29, 2026
📄 Pages 408–420
🔖 ID: IJCT-V13I5P48
Table of Contents
ToggleIncident Knowledge Graphs for Site Reliability Engineering: Connecting Alerts, Runbooks, Services, Deployments, and Postmortems
Author(s)
Venkata Praveen Annam
Abstract
On-call engineers responding to a production incident must typically search across several disconnected systems — alerting dashboards, wikis, service catalogs, deployment logs, and postmortem archives — to reconstruct the operational context needed for diagnosis. This fragmentation slows response and causes institutional knowledge captured in past postmortems to go unused in subsequent, related incidents. This article presents Incident Knowledge Graphs (IKG), a framework that represents alerts, runbooks, services, deployments, and postmortems as typed nodes and relations in a unified, continuously updated property graph, and applies graph traversal combined with dense embedding search to retrieve contextually relevant information during live incidents. The framework was evaluated on a benchmark of 640 held-out on-call retrieval queries and piloted over a twelve-month period across a multi-service production environment. The graph-plus-embedding hybrid retrieval approach achieved 0.86 precision at five results and 0.83 mean reciprocal rank, outperforming keyword search, tag-based lookup, and embedding-only retrieval baselines. Field deployment of the IKG was associated with a reduction in median time to locate a relevant runbook from 9.8 to 1.6 minutes and a reduction in overall median time-to-resolution from 88.0 to 46.0 minutes across 187 tracked incidents. The article presents the graph schema, extraction and construction methodology, retrieval architecture, evaluation results, and discusses the organizational practices that determine whether such a graph remains a living, trustworthy source of institutional memory rather than a stale artifact.
Keywords
site reliability engineering; knowledge graphs; incident response; runbook retrieval; postmortem analysis; institutional knowledge; semantic search; graph databases
Conclusion
This article presented Incident Knowledge Graphs, a framework for unifying alerts, runbooks, services, deployments, and postmortems into a single queryable graph, and applying hybrid graph-and-embedding retrieval to surface contextually relevant knowledge during live incident response. The hybrid retrieval approach outperformed keyword, tag-based, and single-modality baselines on a 640-query benchmark, and its field deployment was associated with substantial reductions in context-retrieval time and a 48 percent reduction in overall median time-to-resolution across 187 tracked incidents. These findings support treating incident-response knowledge not as a collection of disconnected documents to be searched individually, but as a connected graph whose structure itself carries diagnostic value, provided organizations invest in the ongoing content-maintenance practices needed to keep that graph current and trustworthy.
References
Allspaw, J. (2012). Blameless postmortems and a just culture. Etsy Engineering Blog. Beyer, B., Jones, C., Petoff, J., & Murphy, N. R. (Eds.). (2016). Site reliability engineering: How Google runs production systems. O’Reilly Media. Beyer, B., Murphy, N. R., Rensin, D. K., Kawahara, K., & Thorne, S. (Eds.). (2018). The site reliability workbook: Practical ways to implement SRE. O’Reilly Media. Bordes, A., Usunier, N., Garcia-Duran, A., Weston, J., & Yakhnenko, O. (2013). Translating embeddings for modeling multi-relational data. Advances in Neural Information Processing Systems, 26, 2787–2795. Ehsani, M., & Kang, S. (2022). Applications of knowledge graphs in enterprise search: A survey. ACM Computing Surveys, 55(3), 1–36. Hogan, A., Blomqvist, E., Cochez, M., d’Amato, C., Melo, G. D., Gutierrez, C., Kirrane, S., Gayo, J. E. L., Navigli, R., Neumaier, S., Ngomo, A.-C. N., Polleres, A., Rashid, S. M., Rula, A., Schmelzeisen, L., Sequeda, J., Staab, S., & Zimmermann, A. (2021). Knowledge graphs. ACM Computing Surveys, 54(4), 1–37. https://doi.org/10.1145/3447772 Karpathy, R., & Fatemi, B. (2021). Graph-based retrieval augmented generation for domain-specific question answering. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 3915–3926. Karumuri, S., Solleza, F., Zdonik, S., & Tatbul, N. (2021). Towards observability data management at scale. ACM SIGMOD Record, 49(4), 18–23. https://doi.org/10.1145/3456859.3456863 Karypis, K., & Han, E.-H. (2000). Concept indexing: A fast dimensionality reduction algorithm with applications to document retrieval and categorization. Proceedings of the Ninth International Conference on Information and Knowledge Management, 12–19. Karger, D. R., quan Sun, Y., & Rasmussen, C. (2019). Retrieval-augmented question answering with dense passage retrieval. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. Nair, V., Raul, A., Khanduja, S., Bahirwani, V., Shao, Q., Sellamanickam, S., Keerthi, S., Herbert, S., & Dhulipalla, S. (2015). Learning a hierarchical monitoring system for detecting and diagnosing service issues. Proceedings of the 21st ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2029–2038. https://doi.org/10.1145/2783258.2788624 Notaro, P., Cardoso, J., & Gerndt, M. (2021). A survey of AIOps methods for failure management. ACM Transactions on Intelligent Systems and Technology, 12(6), 1–45. https://doi.org/10.1145/3483424 Rooney, S., & Buckley, S. (2020). Learning from incidents: How SRE teams turn failures into resilience. Proceedings of the 2020 USENIX SREcon Conference. Sridharan, C. (2018). Distributed systems observability: A guide to building robust systems. O’Reilly Media.
📋 How to Cite This Paper
Venkata Praveen Annam (2026). Incident Knowledge Graphs for Site Reliability Engineering: Connecting Alerts, Runbooks, Services, Deployments, and Postmortems. International Journal of Computer Techniques, 13(5), 408–420. ISSN: 2394-2231. DOI: https://doi.org/10.5281/zenodo.23040083
Related Posts:










