16 open roles
Manager, Site Reliability Engineering
Job description
- Leadership & Strategy Lead, mentor, and grow the SRE team; set clear goals, on-call structure, and career paths. Define and own SRE strategy, roadmap, and best practices aligned with business and compliance requirements. Drive a culture of reliability, automation, and blameless postmortems.
- Reliability & Availability Own SLAs, SLOs, and SLIs for all production platforms (core banking, APIs, payments). Ensure 99.9%+ availability of critical services and lead efforts to eliminate single points of failure. Manage capacity planning, scalability, and disaster recovery (DR/BCP) strategies.
- Infrastructure & Automation Own and evolve our cloud and on-prem infrastructure (AWS/Azure, Kubernetes, Docker, Terraform). Drive Infrastructure as Code (IaC), CI/CD, and GitOps maturity to enable safe, frequent releases. Lead automation of operational toil, provisioning, and configuration management.
- Incident & Problem Management Own the incident response lifecycle - detection, escalation, resolution, and post-incident review. Build and improve monitoring, alerting, logging, and observability stacks (Prometheus, Grafana, ELK/Datadog, PagerDuty). Act as final escalation for P1/P2 incidents.
- Security & Compliance Partner with Security and Compliance to ensure infrastructure meets PCI-DSS, NDPA, CBN, and ISO 27001 requirements. Embed security, secrets management, and vulnerability remediation into SRE practices. Own change management and audit readiness for infrastructure changes.
- Collaboration Collaborate closely with Software Engineering, Product, Security, and Client Success to ensure reliability is built-in. Provide technical guidance to engineering teams on resilient architecture patterns.
Requirements
Experience 5+ years in DevOps / SRE / Infrastructure Engineering, with 2+ years in a team lead role. Proven experience managing highly available, high-transaction systems in fintech, banking, or large-scale B2B SaaS. Strong track record managing production incidents and on-call teams. Technical Skills Deep expertise in Linux, networking, and distributed systems.
Strong hands-on experience with AWS (EC2, EKS, RDS, VPC, IAM, CloudWatch) or Azure. Expert in Kubernetes, Docker, Terraform, or Ansible. Proficiency in at least one scripting/programming language: Python, Go, or Bash. Experience with CI/CD tools (Jenkins, GitLab CI, GitHub Actions, ArgoCD). Solid understanding of monitoring/observability tools (Prometheus, Grafana, ELK, Datadog, New Relic).
Nice to Have
Experience with core banking systems, payment switches, or ISO 8583. Experience with database reliability (PostgreSQL, MySQL, MongoDB, Redis). Certifications: AWS Solutions Architect / DevOps Engineer, CKA/CKAD. Experience with service mesh (Istio/Linkerd) and chaos engineering. Soft Skills Excellent leadership, communication, and stakeholder management.
Strong analytical and problem-solving mindset under pressure. Ability to balance operational rigor with delivery speed.
Benefits
Qore provides the rare opportunity to make history in the financial space for Africa by Africans, while working with the smartest, brightest & coolest minds in Africa. Our people & culture team continuously thinks of innovative ways to improve employee experience and some of the other benefits of working with Qore includes: Very competitive and rewarding pay Flexible work option (i.
e., Remote work) Paid Lunch for onsite work Lifelong Learnings
Description copied from Qore's careers page. Read the full posting before you apply.
More jobs at Qore
Country Manager, Ethiopia
Qore· Addis AbabaManager, Merchant Services Business Development
Qore· LagosAssociate, Operational Excellence
Qore· LagosEngineering Specialist, Offensive Security
Qore· LagosGraphics Designer
Qore· Lagos
Manager, Site Reliability Engineering jobs at other companies
Manager, Site Reliability Engineering
Mastercard· Toronto, Canada (Ethoca)Manager, Site Reliability Engineering
DriveWealth· Austin, Texas, United States; Dallas, Texas, United States; Miami, Florida, United States; Office - Chicago· $150k – $170kManager, Site Reliability Engineering
M&T Bank· Buffalo, NY· $140k – $233kManager, Site Reliability Engineering
Palo Alto Networks· Office - USA - CA - HeadquartersManager, Site Reliability Engineering
OpenText· Makati City, National Capital Region (NCR), PHL