4,192 open roles
Principal - AI & HPC Data Centre Compute
Job description
We’re looking for a Director – AI & HPC Data Centre Compute to join our team in London, United Kingdom in a hybrid working mode. In this senior leadership role, you will drive EPAM’s AI and HPC data centre strategy, leading engagements that optimise compute infrastructure across full-stack environments—from accelerators and networking to schedulers, containers and AI platforms.
You will address cost, scalability and power constraints through software engineering, platform optimisation and systems integration rather than hardware procurement, helping clients deliver performance and efficiency at scale.
Responsibilities
Advise hyperscalers, neoclouds and enterprises on compute strategies for AI and HPC, including power, cooling and network design Optimise AI training and inference workloads (LLMs, multimodal and scientific AI) across distributed clusters for cost, throughput and latency targets Lead HPC cluster design and orchestration using Slurm, Kubernetes and parallel processing models (MPI) Enhance GPU utilisation by addressing bottlenecks across compute, memory and data pipelines; collaborate with energy teams on power-aware scheduling Architect scalable AI platforms integrating MLOps frameworks, automation tools and reusable assets Develop repeatable AI/HPC offerings, contribute to solution roadmaps and establish strategic partnerships with major ecosystem players Track industry trends in AI factories, GPU economics and liquid cooling technologies, representing EPAM at industry events and thought leadership forums Requirements 12+ years of experience in HPC, AI infrastructure, accelerated computing or distributed systems architecture Deep knowledge of GPU architectures, AI workloads, networking and large-scale cluster operations Hands-on expertise with Slurm, Kubernetes and performance optimisation across multi-node environments Proven ability to advise senior stakeholders and influence technical strategy at C-level Demonstrated experience in pre-sales solutioning and shaping complex technology engagements Nice to have Exposure to NVIDIA ecosystems (DGX, HGX, SuperPOD) or alternative accelerators (AMD or similar) Familiarity with Infini.
Band/RoCE networking, Py. Torch, high-density rack design, liquid cooling or GPU-as-a-Service deployments
Description copied from EPAM Systems's careers page. Read the full posting before you apply.
More jobs at EPAM Systems
Senior Full-Stack Engineer (Python+React)
EPAM Systems· Remote (Serbia)Manager/Senior Manager, Delivery Management - Application Security
EPAM Systems· London, England, UKSenior Data Scientist
EPAM Systems· Remote (Argentina; Brazil; Chile; Colombia; Mexico)Senior End-User Support Engineer
EPAM Systems· Warsaw, Masovian Voivodeship, PolandSenior Functional Tester
EPAM Systems· Remote (Brazil; Argentina; Chile; Colombia; Mexico)
More jobs in London
Public Affairs Manager - 12 month Fixed Term Contract
National Grid· London, GB, WC2N 5EHPrincipal Engineer - Installation and Projects
Subsea 7· London (Sutton), GBSenior Client Account Manager, Global Strategic Accounts (UK)
Reddit· London, United KingdomInsights & Measurement Lead
Overwolf· LondonSales Associate - Clarence Street, Kingston
Skechers· London, United Kingdom