
Job description
- Role Purpose (1–3 lines): Leads complex system administration and compute operations across enterprise server environments, ensuring secure, reliable, scalable and well-governed platforms through advanced troubleshooting, automation, configuration standards, lifecycle management and continuous service improvement. 2.
Key Responsibilities
- Lead end-to-end administration of Linux, Windows and Unix-based compute environments, including provisioning, configuration management, OS tuning, patching, upgrades and lifecycle maintenance.
- Design and implement automation scripts, repeatable workflows and operational runbooks to improve deployment speed, configuration consistency, environment hygiene and manual effort reduction.
- Resolve complex compute platform incidents and problems, lead root-cause analysis, drive corrective actions and provide technical escalation support for high-impact issues.
- Support capacity planning, performance monitoring, resource optimization and compute cost efficiency in partnership with infrastructure, cloud and platform teams.
- Maintain security baseline compliance through vulnerability remediation, system hardening, access control validation and audit-ready operational practices.
- Review technical changes, validate implementation quality, mentor junior engineers and promote engineering standards, documentation discipline and operational maturity.
- Ensure adherence to change control, incident management, access management and regulatory requirements while maintaining accurate configuration records and audit trails.
- Key Performance Indicators (KPIs):
- Compute platform uptime & SLA adherence
- MTTR reduction & escalation avoidance
- Patch compliance / security posture improvement
- Automation coverage & manual effort reduction
- Accuracy & compliance of configuration standards