Call For Paper September 2026

Research Article | Open Access | Download PDF
Volume 13 | Issue 8 | Year 2026 | Article Id. IJCSE-V13I8P102 | DOI : https://doi.org/10.14445/23488387/IJCSE-V13I8P102

A Multi-Agent Generative AI and Machine Learning Framework for Autonomous Incident Intelligence, Predictive Service Assurance, and Customer Experience Optimization in Cloud-Native Digital Service Platforms


Darshan Bhansali

Received Revised Accepted Published
19 Jun 2026 27 Jul 2026 15 Aug 2026 29 Aug 2026

Citation :

Darshan Bhansali, "A Multi-Agent Generative AI and Machine Learning Framework for Autonomous Incident Intelligence, Predictive Service Assurance, and Customer Experience Optimization in Cloud-Native Digital Service Platforms," International Journal of Computer Science and Engineering, vol. 13, no. 8, pp. 11-28, 2026. Crossref, https://doi.org/10.14445/23488387/IJCSE-V13I8P102

Abstract

The growing complexity of the online service platforms based on the cloud has posed new challenges to the traditional IT operations unprecedentedly, and reactive incident management is no longer adequate in ensuring the reliability of service and customer satisfaction. The authors present a Multi-Agent Generative AI and Machine Learning (MAGML) system for autonomous incident intelligence, as well as predictive service assurance and customer experience optimization in this article. The domain-specific ecosystem of intelligent agents to identify root causes, investigate incidents, assess the impact, draw up remediation plans and learn about incidents is all incorporated under the framework incorporation. With a Generative AI-based engine, heterogeneous operational data, such as telemetry, logs, distributed traces, dashboards, and past incidents, are converted into contextual operational intelligence, automated incident summaries, diagnostics, and remediation suggestions. Infrastructure performance is combined with application observability as well as customer journey analytics and business key performance indicators to develop machine learning models that are capable of predicting the occurrence of a service failure prior to the disruption of the service affecting customers. Instead, reinforcement learning enables learning to remediate itself based on outcomes of the learning, and an Operational Knowledge Graph can model the connection between services to increase correlations between events, decrease unnecessary alerts about events, and accelerate the process of root cause detection. Another Customer Experience Intelligence overlay that measures the business impact, the severity of the incident, and SLA/SLO compliance or an AI-based operational copilot that supports the engineer with runbooks, intelligent predictions, and diagnostics. With heterogeneous observability data packets that comprised the HDFS Log Dataset, the BlueGene/L (BGL) Log Dataset, the OpenTelemetry distributed traces, Kubernetes operational metrics and the Application Performance Monitoring (APM) data, the given framework has been experimentally tested and compared with six current approaches, including the threshold-based monitoring, the Random Forest, the LSTM, the Knowledge Graph-based RCA, the RAG-enabled LLM incident intelligence, and the traditional enterprise AIOps platform. Laboratory results demonstrate that the proposed MAGML model achieves an accuracy, precision, recall and F1-score of 98.4 per cent, 98.1 per cent, 98.3 per cent, and 98.2 per cent, respectively, and reduces MTTD to 3.1 minutes and MTTR to 14.6 minutes. It also boasts of 98.6% accuracy RCA, 98.3% accuracy alert-correlation, 99.1% SLA compliance and 99.0% recoverability success, which confirms that it is solid in its predictability incident management model, autonomous remediation model, service assurance model, and customer-focused operation insight of cloud-native digital service.

Keywords

Multi-Agent Systems; Generative AI; AIOps; Predictive Incident Management; Cloud-Native Microservices; Customer Experience Intelligence; Reinforcement Learning.

References

  1. Juncal Alonso, Leire Orue-Echevarria, and Maider Huarte, “CloudOps: Towards the Operationalization of the Cloud Continuum: Concepts, Challenges and A Reference Framework,” Applied Sciences, vol. 12, no. 9, pp. 1-24, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  2. Xinyi Hou et al., “Large Language Models for Software Engineering: A Systematic Literature Review,” ACM Transactions on Software Engineering and Methodology, vol. 33, no. 8, pp. 1-79, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  3. Patick Lewis et al., “Retrieval-augmented Generation for Knowledge-Intensive NLP Tasks,” Advances in Neural Information Processing Systems, vol. 33, pp. 1-16, 2020.
    [
    Google Scholar] [Publisher Link]
  4. Shanshan Han et al., “LLM Multi-Agent Systems: Challenges and Open Problems,” arXiv preprint, pp. 1-8, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  5. Qingyun Wu et al., “AutoGen: Enabling Next-gen LLM Applications Via Multi-Agent Conversation,” arXiv preprint, pp. 1-43, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  6. Silverio Martínez-Fernández et al., “Software Engineering for AI-based Systems: A Survey,” ACM Transactions on Software Engineering and Methodology, vol. 31, no. 2, pp. 1-59, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  7. D. Russo, S. Van Berkel Baltes, and Christop Treude, “Generative AI in Software Engineering must be Human-Centered: The Copenhagen Manifesto,” Journal of Systems and Software, vol. 216, no. 1, pp. 1-2, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  8. Yuxuan Jiang et al., “XPERT: Empowering Incident Management with Query Recommendations Via Large Language Models,” Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (ICSE), Lisbon, Portugal, pp. 1-13, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  9. Yinfang Chen et al., “Automatic Root Cause Analysis Via Large Language Models for Cloud Incidents,” Proceedings of the Nineteenth European Conference on Computer Systems, Athens, Greece, pp. 674-688, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  10. Pengxiang Jin et al., “Assess and Summarize: Improve Outage Understanding with Large Language Models,” Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE), San Francisco, USA, pp. 1657-1668, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  11. Yinfang Chen et al., “AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds,” Proceedings of Machine Learning and Systems, vol. 7, 2025.
    [
    Google Scholar] [Publisher Link]
  12. Manish Shetty et al., “Building AI Agents for Autonomous Clouds: Challenges and Design Principles,” Proceedings of the 2024 ACM Symposium on Cloud Computing, Redmond, USA, pp. 99-110, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  13. Zhaoyang Yu et al., “MonitorAssistant: Simplifying Cloud Service Monitoring Via Large Language Models,” Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, Porto de Galinhas, Brazil, pp. 38-49, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  14. Pradeep Dogga et al., “AutoARTS: Taxonomy, Insights and Tools for Root Cause Labelling of Incidents in Microsoft Azure,” 2023 USENIX Annual Technical Conference (USENIX ATC 23), pp. 359-372, 2023.
    [
    Google Scholar] [Publisher Link]
  15. Supriyo Ghosh et al., “How to Fight Production Incidents? An Empirical Study on a Large-scale Cloud Service,” Proceedings of the 13th Symposium on Cloud Computing, San Francisco, California, pp. 126-141, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]