astra-north

Site Reliability Engineer (SRE) – Dynatrace & AI Observability

astra-north

Toronto, Ontario, CanadaContractPosted Jul 6, 2026

Job description

Site Reliability Engineer Location: Toronto, ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required Skills

  • Site Reliability Engineering (SRE)
  • DevOps
  • Dynatrace Role Summary
  • Design, implement, and optimize Site Reliability Engineering (SRE) and DevOps practices to ensure high system availability, performance, and reliability across distributed environments.
  • Leverage Dynatrace, Davis AI, automation, and cloud technologies to enable proactive monitoring, intelligent automation, and operational excellence. Role Description Dynatrace & AI-Driven Observability
  • Lead the implementation and optimization of the Dynatrace platform across applications and infrastructure.
  • Utilize Dynatrace Davis AI for automated root cause analysis, anomaly detection, event correlation, predictive performance insights, and alert noise reduction.
  • Configure and manage OneAgent deployments, Smartscape topology mapping, service flow, and distributed tracing.
  • Define and monitor SLIs, SLOs, and user experience metrics.
  • Build custom dashboards, alerts, and observability pipelines.
  • Integrate Dynatrace with CI/CD pipelines for release validation and performance gating.
  • Integrate Dynatrace with incident management tools such as PagerDuty and ServiceNow.
  • Enable self-healing automation using Dynatrace event triggers and AI-driven insights. Automation & Configuration Management
  • Design and implement automation solutions using Ansible.
  • Automate configuration management, application deployments, and environment provisioning.
  • Develop reusable Ansible playbooks and roles for scalable operations.
  • Automate operational tasks, patching, compliance processes, and remediation workflows.
  • Integrate Ansible with CI/CD pipelines and monitoring systems. Cloud & DevOps
  • Design and manage cloud-native solutions on AWS with exposure to Azure.
  • Develop infrastructure using Terraform, CloudFormation, or AWS CDK.
  • Build and manage CI/CD pipelines using GitHub Actions, Jenkins, or GitLab CI.
  • Develop and deploy serverless solutions using AWS Lambda, API Gateway, and Step Functions.
  • Automate DevOps and operational workflows using Python (boto3) and Bash scripting.
  • Deploy and maintain production environments through automated pipelines.
  • Optimize cloud infrastructure for cost, performance, and scalability. Monitoring & Reliability Engineering
  • Monitor and manage AWS CloudWatch and Azure Monitor/Log Analytics.
  • Design unified observability across multi-cloud environments.
  • Implement logging and distributed tracing strategies.
  • Work with Docker, Kubernetes, ECS, and AKS environments.
  • Design fault-tolerant, highly available, and disaster recovery solutions.
  • Support incident response, on-call activities, and root cause analysis (RCA). Required Qualifications
  • Proven experience with Dynatrace APM, Real User Monitoring (RUM), and infrastructure monitoring.
  • Strong hands-on experience with Dynatrace Davis AI capabilities.
  • Experience with Ansible for automation and configuration management.
  • Deep knowledge of AWS services and cloud-native architectures.
  • Experience with Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.
  • Proficiency in Python (boto3) and Bash scripting.
  • Experience supporting production-scale environments.
  • Business Analyst experience.
  • Scrum Master experience.

Nice to Have

  • Dynatrace Associate or Professional certification.
  • Experience with Dynatrace APIs and automation.
  • Experience building self-healing systems using AI-driven triggers.
  • Familiarity with Prometheus, Grafana, and the ELK Stack.
  • Azure cloud experience and certifications.
  • Experience with GitOps and Platform Engineering.