Slide logo
Slide

8 open roles

Staff Cloud Operations Engineer

Remote - USFull-timePosted Oct 9, 2026

Job description

ABOUT SLIDE Slide is a modern, security-first Business Continuity & Disaster Recovery (BCDR) company built exclusively for Managed Service Providers. Founded by Austin McChord (Datto Founder & former CEO) and Michael Fass (former Datto General Counsel & Chief People Officer), Slide is led by a team of industry veterans with deep expertise in backup, disaster recovery, and cybersecurity.

Built from scratch, from a clean-room code base, free from legacy technical debt, to deliver the MSP-centric backup and recovery platform of the future. By focusing on security, performance, and simplicity, Slide provides a powerful, cost-effective, and easy-to-use solution that ensures MSPs can protect their clients' data without the constraints of outdated technology and restrictive pricing models.

About the Job

As a Staff Cloud Operations Engineer at Slide, you'll design and run the systems our product and cloud depend on. That includes the Terraform and Ansible platforms we build on, the pipelines that ship code to production, and the data center fleet that powers our products. You'll make changes safely through good review, guarded applies, config in code and alerts that matter.

When the same operational problem keeps coming back, you'll fix it for good with a platform or a playbook. This is a hands-on Staff role on a small team. Leadership here means setting the architecture and standards, and making every engineer who comes after you faster.

Responsibilities

What you'll work on

  • Product platforms. Build declarative cloud foundations (cells and zones, data stores, identity, images) so Engineering can ship without waiting on custom infrastructure each time.
  • Faster, safer delivery. Make the path from code to image to production automated, reviewable and safe: Packer and CI image builds, Terraform with gated applies, merge queues and branch protection, and predictable deploys of tagged releases.
  • Self-service. Turn repeated requests (DNS, access, repo and environment scaffolding, certificates) into infrastructure-as-code and CI/CD, so Infra multiplies the team instead of becoming its queue.
  • Data center and cloud storage fleet. Run the fleet behind Slide's cloud storage as we add sites and carriers. Lifecycle, networking, storage and observability should run on playbooks, not heroics.
  • AI-driven operations. We're a small team by design and an AI-forward company. You'll use AI to speed up how we build, operate and troubleshoot infrastructure, from writing and reviewing IaC to triaging alerts and automating runbooks, so a few engineers can do the work of many. How you'll work
  • Own architecture decisions for our data center and cloud infrastructure, and help shape the roadmap.
  • Build the monitoring, logging and alerting that let us find problems before customers do.
  • Lead the hardest incidents and the follow-up work that keeps them from coming back.
  • Review infrastructure code, set the standards and mentor the engineers around you.
  • Work alongside engineers and peer teams across the company. We're growing fast, so you'll pitch in wherever infrastructure can unblock the work.

Qualifications

  • Experience. 5-10 years in infrastructure, cloud operations, SRE or a related role, including time operating production systems at scale.
  • Infrastructure-as-code in production. You've shipped Terraform or Ansible and operated what you built.
  • Delivery pipelines for infrastructure. You know how images and artifacts, CI, gated applies and deploys, and dev, staging and production environments fit together.
  • Deep Linux. When the abstraction leaks, you can work on real hosts over SSH, serial console or out-of-band management.
  • Cloud platforms. You have hands-on experience with public clouds (GCP, AWS, etc.), including networking, IAM, managed data stores and compute fleets.
  • Observability. You connect metrics to alerts to runbooks, and you build alerts that point to where the failure actually is.
  • Clear writing. You work through pull requests and well-documented systems, and you're comfortable collaborating asynchronously on a hybrid team.
  • AI in your workflow. You already use AI tools to work faster, and you're eager to bring them into how we run infrastructure.
  • On-call. You'll join the on-call rotation and occasionally coordinate data center visits or remote hands.

Benefits

  • Comprehensive health, dental, and vision coverage.
  • Paid Time Off: Generous paid time off and holiday schedule.
  • Retirement Plan: 401(k)
  • Professional Development: Opportunities for training and professional growth.
  • The opportunity to do your life's work in a dynamic and creative environment with like-minded individuals.

Description copied from Slide's careers page. Read the full posting before you apply.

More jobs at Slide

See all openings at Slide