astra-north

Site Reliability Engineer – APM, Dynatrace, Observability

astra-north

Toronto, Ontario, CanadaContractPosted Jun 17, 2026

Job description

Location: Toronto, ON (Hybrid – 4 Days Onsite) Duration: 12 Months Experience: 6–8 Years Required Skills Kubernetes & Containers (Kubernetes 1.24+, CRDs, Operators) Prometheus, Grafana, Thanos/Cortex/Mimir, Alertmanager Loki/ELK/Splunk, Open. Telemetry, Jaeger/Tempo Git. Ops (ArgoCD/FluxCD), Helm, Terraform AWS/Azure/GCP, S3/Object Storage PromQL, LogQL/Lucene, SQL Python, Bash or Go CI/CD (GitHub Actions, GitLab CI, Jenkins) Key Responsibilities Design, deploy, and manage enterprise observability platforms across Kubernetes environments.

Build and maintain monitoring, logging, tracing, dashboards, and alerting solutions. Automate deployments using Git. Ops and Infrastructure as Code. Optimize observability performance, scalability, and cost. Integrate observability with CI/CD pipelines, cloud platforms, and incident management tools. Collaborate with SRE, DevOps, and application teams to improve platform reliability and reduce MTTR.

Preferred Skills

AI/ML for Observability, AIOps, LLM integration Kubernetes Certifications (CKA/CKAD/CKS) Grafana Certification Experience with SLIs, SLOs, Error Budgets Financial Services or regulated industry experience APM tools and Service Mesh observability