Deputy Incharge - Engineering Delivery & Reliability
National Payments Corporation of India
Hyderabad, TelanganaFull TimePosted Aug 3, 2026
Job description
About the Role
We are seeking a Incharge/Deputy Incharge – Site Reliability Engineering to lead and scale our on-premises payments infrastructure. This role requires deep expertise in hybrid environments (bare metal, virtual machines, and containerized workloads) along with a strong focus on system reliability, performance, and security.
You will drive platform resilience strategy, lead high-performing SRE teams, and integrate AI/ML-driven automation and observability into infrastructure operations. Location: Hyderabad Experience: 15+ years Key Responsibilities Infrastructure & Platform Engineering
- Lead the design, operation, and scaling of on-premise infrastructure supporting mission-critical payments systems
- Manage hybrid environments: o Virtual machines (VMware/KVM) o Containerized workloads (Docker on bare metal, Kubernetes – preferred)
- Ensure high availability, fault tolerance, and disaster recovery readiness Middleware & Distributed Systems
- Oversee and optimize: o Nginx (reverse proxy, traffic routing, load balancing) o Redis (caching, clustering, HA, persistence tuning) o Kafka (streaming, partition design, replication, performance tuning)
- Drive architecture improvements for low latency and high throughput systems IT Ops
- Build and scale CI/CD pipelines using Jenkins (or equivalent tools)
- Implement Git. Ops and infrastructure-as-code practices
- Lead automation initiatives using: o Ansible o Shell/Python scripting SRE & Reliability Engineering
- Define and enforce SLOs, SLIs, and SLAs
- Lead incident management, RCA, and postmortem culture
- Drive proactive monitoring, alerting, and observability strategy Networking & Security
- Manage and troubleshoot: o TCP/IP, DNS, Load Balancing o Firewalls, WAF, CDN (Akamai preferred)
- Ensure security, compliance (PCI-DSS preferred), and governance standards
- Lead vulnerability management and risk mitigation AI-Driven Operations & Innovation
- Introduce AI/ML-based observability, anomaly detection, and predictive maintenance
- Leverage AI tools for: o Incident prediction and auto-remediation o Intelligent log analysis o Capacity planning and forecasting
- Evaluate and integrate AIOps platforms Leadership & Stakeholder Management
- Lead and mentor a team of SREs/DevOps engineers (minimum 3+ years of people management)
- Drive hiring, performance management, and capability building
- Collaborate with Engineering, Security, Product, and Infra teams
- Own strategic roadmap for platform scalability and resilience Governance & Operations
- Maintain infrastructure inventory, asset lifecycle, and documentation
- Ensure adherence to audit and compliance requirements
- Drive operational excellence and cost optimization ________________________________________ Required Qualifications
- 15+ years of experience in Infrastructure, SRE, or DevOps roles
- 3+ years in leadership/people management roles
- Strong experience managing on-premise or hybrid infrastructure at scale
- Deep hands-on expertise in: o Linux (RHEL/CentOS/Ubuntu) o Nginx, Redis, Kafka o Jenkins, Git-based workflows o Ansible and scripting
- Strong understanding of networking and security concepts
- Proven experience in high-availability systems and production support ________________________________________ Preferred Qualifications
- Exposure to Kubernetes and container orchestration platforms
- Experience in payments domain or financial services (highly desirable)
- Familiarity with PCI-DSS or similar compliance frameworks
- Experience with AIOps tools (e.g., Grafana, Victoria Logs, Victoria Metrics)
- Knowledge of cloud integration (AWS/Azure hybrid setups)
- Strong programming skills (Python/Go – good to have)