Job description
Systems fail. That's not a risk to eliminate — it's a certainty to prepare for. The real question is whether we already know how fast we recover when it happens, and who finds out first: us, or our traders. You'll be the person who makes sure it's always us. Why This Matters Deriv's mission is Trading for Anyone, Anywhere, Anytime.
That promise doesn't hold if a database failure, a regional outage, or a SaaS provider going dark takes us offline while traders are mid-position. Resilience isn't a compliance checkbox here — it's the thing standing between "we recovered in nine minutes" and "we lost customer trust for a quarter." We're already building AI into how we detect and respond to failure: autonomous monitoring that correlates alerts against historical patterns, remediation guidance generated in plain language instead of hours of manual digging, and automated evidence trails that used to take a team days to assemble by hand.
DR is the next frontier for that same thinking — and you'll be the one applying it. Why Deriv We're already in production, not planning. AI-driven alert correlation and remediation guidance running today, not on a roadmap Automated evidence-gathering workflows replacing manual audit prep across the business A culture that treats "we tested it and it broke" as useful information, not a failure We share what we learn.
Deriv is where we write about what we're building, what breaks, and what we figure out the hard way. You'll join a resilience function that's actively being rebuilt around AI tooling — not one waiting for someone else to modernise it first. What You’ll Do Run real recovery, not just recovery plans Design and execute DR exercises for business-critical applications — database failures, regional outages, SaaS provider disruptions, live failover — and report honestly on what held up and what didn't Work hands-on with DevOps, WinOps, SRE, engineering, product, operations, and SaaS owners to test recovery plans against how systems actually behave under stress Feed every incident and near-miss straight back into the DR plan, so the same failure never catches us twice.
Make the numbers mean something Conduct business impact analyses that pressure-test RTO and RPO targets against real operational data, not last year's assumptions Push back when a recovery target looks good on paper but wouldn't survive contact with a real outage Keep the evidence honest Maintain and audit DR strategies, runbooks, test records, BIAs, and recovery evidence so they're accurate today, not just accurate when they were written Build dashboards that give leadership a real answer to "are we actually ready?
" — not a static slide from last quarter Use AI as your first move, not your last resort Model failure scenarios, generate exercise briefs, and simulate impact using AI tooling instead of building everything from a blank template Automate evidence gathering and gap surfacing so your time goes to the failures that need judgment, not paperwork Analyse historical incidents for patterns humans tend to miss under deadline pressure Speak up before it's a postmortem Flag it when a new system goes into production without a tested recovery plan — even when nobody asked you to check Translate what's actually at risk into language that works for an infrastructure engineer and for a C-suite leader, in the same conversation if you have to Who You Are You solve problems independently, and you don't hoard the answer.
You handle routine DR gaps and exercise findings without waiting to be told what to do next — and when you figure something out, you make sure the rest of the team doesn't have to learn it the hard way too. You go looking for the failure, not just the checklist Tabletop discussions are a start, not a finish line. You've run real failover tests where something actually broke, and you know the difference between a plan that reads well and a plan that survives contact with reality.
You use AI like a colleague who never sleeps You reach for AI tooling to model scenarios, draft exercise briefs, and chase down evidence gaps — not because it's expected of you, but because it's obviously the faster, better way to work. You write like the audit is tomorrow Your runbooks, BIAs, and recovery narratives are precise enough that someone else could execute them without you in the room, and current enough that they'd actually work if they had to.
You explain risk without dumbing it down or drowning people in it You can tell an engineer exactly what broke and tell a director exactly what it means for the business — and you know which one the room in front of you needs.
Requirements
4+ years in disaster recovery, business continuity, infrastructure resilience, or cloud operations Hands-on experience with AWS, GCP, Azure, or equivalent, including cloud-native failure modes Practical experience running real DR exercises or failover tests — not only writing plans or facilitating tabletops Experience conducting or contributing to business impact analyses, especially in cloud or hybrid environments Working knowledge of ITIL and business continuity frameworks, applied in practice Comfort using AI tools for analysis, reporting, simulation, and evidence automation A relevant certification held or actively in progress — CBCP, MBCI, AWS/GCP Professional, or equivalent Bonus Points Preparing DR evidence for Compliance, Risk, or external audit reviews Building DR readiness dashboards or AI-assisted resilience reporting Integrating AI-driven monitoring or anomaly detection into infrastructure workflows The Honest Reality Most of the time, nothing breaks — and that can feel like nothing's happening.
This role rewards people who stay sharp when the system's quiet, because the exercise you skip is the failure mode you'll meet for real, at the worst possible time. You'll also spend real energy getting busy engineering teams to prioritise a test for a disaster that (hopefully) never comes. That's a harder sell than it sounds, and you'll need the credibility and the evidence to make it stick.
If you want a role where you document plans and hope someone else stress-tests them, this isn't it. If you want to be the reason the outage was a non-event instead of a headline, it might be.