Monitoring & Observability.
Observability engineering — metrics, logs, traces, Prometheus, Grafana, OpenTelemetry, and alerting strategies.
Beginner
Start here — no prior experience needed
Monitoring & Observability Essentials
Learn the difference between monitoring and observability, core signals (metrics/logs/traces), and how to think in SLOs and actionable alerts.
Alerting & Runbooks
Design alerts that page the right people at the right time—plus how to attach runbooks so incidents are fixable quickly.
Intermediate
For developers with core concepts down
Advanced
Production-grade patterns for experienced engineers
OpenTelemetry: Unified Observability
Instrument applications with OpenTelemetry to emit traces, metrics, and logs in a vendor-neutral format. Learn the OTel data model, collectors, and backends.
Incident Management & Post-Mortems
Build a repeatable incident response process: detection, communication, mitigation, and blameless post-mortems that improve reliability over time.