astra-north

Site Reliability Engineer (SRE) – Azure AKS, Observability & Terraform.

astra-north

Montreal, Quebec, CanadaFull timePosted May 4, 2026

Job description

Site Reliability Engineer (SRE) – Azure AKS, Observability & Terraform Key Responsibilities Observability, SRE, DevOps roles with expertise in infrastructure and application reliability Dynatrace, ELK, Splunk, Pager. Duty SLI/SLO frameworks Azure Kubernetes Service (AKS), Terraform, Azure managed services What will you do Design and implement observability-as-code solutions using Terraform for monitoring pipelines, dashboards, and alerting across distributed systems Drive observability improvements using Dynatrace, ELK, Splunk, Pager.

Duty for real-time performance insights and system visibility Instrument applications for end-to-end observability including distributed tracing, metrics collection, and log aggregation across Node.js and .NET microservices and event-driven architectures Troubleshoot complex production incidents across service layers, databases, caches, and APIs using SLI/SLO frameworks Investigate and resolve Azure Kubernetes Service (AKS) infrastructure issues ensuring reliability and scalability of containerized workloads using Terraform and Azure services (SQL MI, Redis, Functions, Event Grid) Translate business requirements into observable, resilient systems aligned to SLIs/SLOs Automate operational tasks using Infrastructure-as-Code and CI/CD to reduce toil and improve resilience Lead incident response and remediation for critical systems, including blameless postmortems and chaos engineering practices Collaborate with development, platform, and business teams to improve availability, scalability, and operational excellence What do you need to succeed Must-have 8+ years experience in SRE, DevOps, or Observability roles focused on infrastructure and application reliability Strong expertise in Dynatrace, ELK, Splunk, Pager.

Duty and observability principles (instrumentation, correlation IDs, SLIs/SLOs) Advanced proficiency in Azure Kubernetes Service (AKS), Terraform, and Azure managed services (SQL MI, Redis, Functions, Event Grid) Hands-on experience with observability instrumentation (distributed tracing, metrics, logs) across Node.js and .

NET microservices and event-driven systems Strong troubleshooting skills across distributed systems (services, databases, caches, APIs) in production environments Incident management expertise using Pager. Duty and Service. Now, including high-severity incident resolution and RCA Knowledge of incident, problem, and change management, SRE principles, blameless postmortems, and chaos engineering Strong communication and leadership skills for cross-functional coordination and incident handling