Société Générale

Site Reliability Engineer (SRE) / Production Engineer

Société Générale

Casablanca, MoroccoPermanent contractPosted Mar 24, 2026

Job description

As a Site Reliability Engineer / Production Engineer, you will be responsible for the reliability, availability, security, and operational excellence of critical platforms operating in a hybrid cloud environment.

You will act as a production owner and infrastructure reliability expert, ensuring smooth operations, secure systems, efficient deployments, and successful large‑scale transformation initiatives.

You work closely with engineering, infrastructure, security, and program teams to ensure systems are robust, secure, and continuously deliver value with minimal downtime.

Key Responsibilities

  • Reliability & Availability
  • Ensure high service availability and business continuity
  • Measure and manage reliability using SLIs, SLOs, and error budgets
  • Design, operate, and continuously improve resilient hybrid cloud architectures

Incident & Production Management

  • Own production incident management end‑to‑end, focusing on MTTR reduction
  • Lead and coordinate major incidents and production outages
  • Define and maintain on‑call rotations and escalation processes
  • Facilitate post‑incident reviews (postmortems) and track remediation actions
  • Exercise authority to block or delay releases when production risk is unacceptable

Monitoring & Observability

  • Build effective, actionable monitoring and alerting:

  • Clear ownership

  • Strong root‑cause signals

  • Defined next steps (runbooks or automation)

  • Implement capacity planning and proactive failure detection

Automation & Operational Excellence

  • Automate operational tasks (IaC, runbooks, self‑healing mechanisms)
  • Improve production readiness and operational maturity
  • Proactively reduce incident risk through continuous improvements

Deployment Pipelines & Release Management

  • Manage and maintain deployment pipelines (CI/CD) for critical applications
  • Ensure reliable, timely, and repeatable application deployments
  • Optimize deployment strategies to minimize downtime and production risk (e.g., rolling deployments, blue/green, canary releases)
  • Improve deployment automation, consistency, and traceability
  • Collaborate with engineering teams to align delivery pipelines with SRE and production standards

Infrastructure Management

  • Contribute to the operation and evolution of infrastructure across cloud and on‑prem environments
  • Ensure stability, performance, scalability, and maintainability of platforms

Drive standardization and industrialization of infrastructure components

Support for Transformation Projects

  • Actively support technology transformation initiatives (cloud adoption, modernization, migration programs)
  • Engage early in projects to secure architectural and operational choices
  • Ensure controlled, reliable go‑lives for new systems and platforms
  • Challenge solutions that do not meet production, reliability, or scalability standards

Security Operations & Vulnerability Management

  • Execute and track security recommendations and remediation actions
  • Manage server and infrastructure patching
  • Identify, assess, prioritize, and remediate vulnerabilities
  • Work closely with security teams to reduce operational and security risk

Your Mindset

  • Strong sense of production ownership and accountability
  • Reliability‑first and security‑aware mindset
  • Ability to balance delivery speed with system stability
  • Comfortable working in cross‑functional, collaborative environments