Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps Using Open Source Tools
- Understanding AIOps concepts and their advantages
- The role of Prometheus and Grafana in the observability ecosystem
- The place of ML in AIOps: comparing predictive and reactive analytics
Configuring Prometheus and Grafana
- Installing and tuning Prometheus for time series data collection
- Building Grafana dashboards with real-time metrics
- Working with exporters, relabeling, and service discovery
Preparing Data for Machine Learning
- Extracting and processing Prometheus metrics
- Preparing datasets for anomaly detection and forecasting tasks
- Leveraging Grafana transformations or Python-based pipelines
Leveraging Machine Learning for Anomaly Detection
- Introduction to ML models for outlier detection (e.g., Isolation Forest, One-Class SVM)
- Training and assessing models on time series data
- Displaying anomalies within Grafana dashboards
Predicting Metrics with Machine Learning
- Developing basic forecasting models (intro to ARIMA, Prophet, LSTM)
- Anticipating system load and resource consumption
- Utilizing predictions for proactive alerting and scaling decisions
Combining ML with Alerting and Automation
- Establishing alert rules based on ML outputs or specific thresholds
- Configuring Alertmanager and notification routing
- Initiating scripts or automation workflows upon anomaly detection
Scaling and Operationalizing AIOps
- Integrating with external observability solutions (e.g., ELK stack, Moogsoft, Dynatrace)
- Deploying ML models within observability pipelines
- Best practices for implementing AIOps at scale
Recap and Future Directions
Requirements
- A solid grasp of system monitoring and observability principles
- Prior experience with Grafana or Prometheus
- Knowledge of Python and fundamental machine learning concepts
Target Audience
- Observability Engineers
- Infrastructure and DevOps Teams
- Monitoring Platform Architects and Site Reliability Engineers (SREs)