
Senior Site Reliability Engineer — CDN Infrastructure
Job description
About the Role
We're looking for a Senior SRE with 5+ years of experience to help operate and evolve the infrastructure behind our CDN — from the baremetal edge nodes serving traffic, to the Kubernetes-based control plane, to the observability and data pipelines that tell us what's actually happening on the network. You'll work across the full stack: provisioning and configuration management, CI/CD, Kubernetes, and the logging/monitoring systems that let us catch problems before customers do.
This is a hands-on, senior individual-contributor role for someone who's comfortable moving between low-level networking debugging and higher-level platform/infra automation, and who can drive decisions with less oversight.
What You'll Do
Operate and maintain our fleet of self-managed baremetal edge servers, using Salt. Stack (or Ansible) for configuration management and automation Manage and improve our managed Kubernetes cluster, deploying and maintaining services via Helm Build and maintain CI/CD pipelines (GitHub Actions) for infrastructure and service deployments Own and extend observability tooling — Prometheus, Grafana, Alertmanager, and distributed tracing with Grafana Tempo Maintain and query Click.
House at scale — millions of rows ingested daily, with a focus on schema design and low-latency query tuning Manage cloud service configuration as code using Terraform Diagnose and resolve issues across the stack — from DNS resolution and BGP/routing anomalies to TCP/IP-level performance regressions and application.
What You'll Need
Requirements What You'll Need 5+ years of experience in an SRE, DevOps, or infrastructure/platform engineering role Deep understanding of networking fundamentals — DNS, TCP/IP, and CDN concepts (caching, routing, anycast, edge delivery) — you should be comfortable reading a packet capture or debugging a DNS resolution chain Solid experience operating Linux systems in production, including self-managed baremetal infrastructure Hands-on experience with configuration management tools — Salt.
Stack or Ansible Experience with Kubernetes in production, including deploying and managing services via Helm Experience building CI/CD pipelines, ideally with GitHub Actions Working knowledge of Terraform or similar IaC tools Practical experience with Prometheus, Grafana, and Alertmanager for monitoring and alerting; familiarity with distributed tracing (Grafana Tempo or similar) Ability to write code in Go or Python for automation, tooling, or internal services Strong debugging skills across layers — from kernel/network to application to infrastructure automation Good communication skills and comfort working in a small, high-ownership team Nice to Have Experience with BGP / anycast routing in a production CDN or network operator context Experience with Click.
House or another columnar/analytical database at scale Experience with Kafka/Redpanda or similar streaming systems for log/data pipelines Prior experience at a CDN, ISP, hosting provider, or similar network-heavy operator Experience with Git. Ops workflows (ArgoCD or similar)