Agentic AIOps for Cloud Log Anomaly Triage with Retrieval-Grounded Root-Cause Explanations

Authors

  • Peng Liu computer science, columbia, ny, usa Author
  • Jessica White computer science, umich, mi, usa Author
  • Wei Han computer science, purdue, in, usa Author
  • Christopher Scott computer science, uci, ca, usa  Author

DOI:

https://doi.org/10.63575/CIA.2024.20215

Keywords:

AIOps, system logs, anomaly detection, root-cause analysis, retrieval-augmented generation, incident triage, OpenStack, Hadoop

Abstract

Cloud operations teams must distinguish genuine failures from benign log variation and then explain the likely root cause with evidence that an operator can inspect. This paper presents a retrieval-grounded AIOps triage agent that converts raw OpenStack and Hadoop logs into normalized event templates, combines lexical, structural, and timing anomaly signals, retrieves analogous incidents, ranks candidate causes, and verifies every explanation against direct log cues or matching historical cases. The evaluation uses all 55 labeled Hadoop application runs and a held-out OpenStack stream containing 198 VM sessions, including the four VM identifiers marked as anomalous by the corpus. Hadoop anomaly detection is measured with three repetitions of stratified five-fold cross-validation; OpenStack models are fitted on normal1, calibrated on normal2, and tested on the abnormal stream. On Hadoop, the full fusion obtains precision 0.913, recall 0.871, F1 0.891, false-positive rate 0.333, and area under the precision–recall curve 0.951. A numeric random forest reaches a similar F1 of 0.894 but raises the false-positive rate to 0.576. On OpenStack, the calibrated fusion detects all four anomalous VMs and produces one false alarm among 194 normal VMs, yielding precision 0.800, recall 1.000, and F1 0.889. For Hadoop root-cause analysis, the verified ranker attains Top-1 accuracy 0.818, Top-2 accuracy 1.000, mean reciprocal rank 0.909, and NDCG@3 0.933 across machine-down, network-disconnection, and disk-full incidents. Verification raises evidence precision from 0.906 to 1.000 while preserving cause accuracy. The results show that retrieval is most valuable for ranked diagnosis and evidence formation, whereas explicit latency modeling is indispensable for the injected OpenStack anomalies.

Author Biography

  • Christopher Scott, computer science, uci, ca, usa 

     

     

     

Published

2024-08-14

How to Cite

[1]
Peng Liu, Jessica White, Wei Han, and Christopher Scott, “Agentic AIOps for Cloud Log Anomaly Triage with Retrieval-Grounded Root-Cause Explanations”, Journal of Computing Innovations and Applications, vol. 2, no. 2, pp. 162–176, Aug. 2024, doi: 10.63575/CIA.2024.20215.