4,198 open roles
Platform Engineering Team Lead
Job description
We are seeking an experienced and visionary Platform Engineering Team Lead to lead a team of highly skilled platform experts. In this role, you will own the architecture, stability, and scaling of our enterprise data streaming and processing platforms. You will act as a key technical leader, bridging the gap between Infrastructure, Data Engineering, and Data Science, ensuring high availability, continuous automation, and state-of-the-art platform observability.
This opportunity is with Naya, an EPAM company, a leading global provider of data platforms and development professional services. Based in Israel, Naya is one of the fastest-growing companies in the data and development technology space, and we’re growing our team.
Responsibilities
Provide professional and personal management, mentorship, and guidance to a team of platform engineering experts Confluent/Kafka Platform Ownership: End-to-end responsibility for the Confluent Suite, including Apache Kafka, Kafka Connect, Schema Registry, REST Proxy, Confluent Cloud, and KSQL Data Platforms: Own and manage enterprise data analytics platforms, including Azure Synapse Data.
Ops, MLOps, & Compute: Manage end-to-end infrastructure supporting Data. Ops, MLOps, and specialized GPU compute resources Ingestion & Logging Infrastructure: Oversee streaming ingestion and logging pipelines based on Apache Ni. Fi, PortX, Fluent Bit, and Filebeat Lifecycle & Projects: Lead complex architecture, installation, upgrading, and migration projects for all data platforms Incident Management: Drive deep troubleshooting and resolution of complex production issues Architectural Guidance: Provide professional architectural advisory and support to Data Engineering and Data Science development teams Vendor Management: Interface directly with external vendors and technology partners Automation & Observability: Drive automation initiatives and implement advanced system monitoring and observability frameworks System Resiliency: Lead initiatives for Capacity Planning, High Availability (HA), and Disaster Recovery (DR) Requirements Leadership: At least 3 years of experience leading and managing a technology/engineering team Domain Expertise: At least 5 years of experience in provisioning, maintaining, and operating large-scale Data and/or Streaming platforms Operating Systems: Deep, hands-on experience working in Linux environments Kafka & Confluent: Significant hands-on experience with on-premises Apache Kafka / Confluent and all its core ecosystem components Production Operations: Proven track record in troubleshooting and root-cause analysis of complex issues in high-pressure Production environments Cross-Functional Collaboration: Extensive experience working closely with Software/Data Developers and Solutions Architects Big Data Ecosystem: Strong hands-on experience with at least one or more of the following: Hadoop, Cloudera, Databricks, Apache Spark, or Azure Synapse Configuration Management: Solid experience working with Ansible for automation Modern Paradigms: Hands-on experience in at least one of these domains: Data.
Ops, MLOps, or AI/ML Platforms A systemic, holistic approach to system architecture and planning Excellent self-learning capabilities and adaptability to new technologies Outstanding interpersonal skills with a strong service-oriented mindset Proven ability to work effectively across multiple cross-functional departments (interfaces) Nice to have Data Ingestion: Strong hands-on experience with Apache Ni.
Fi (Highly Advantageous) Cloud Infrastructure: Experience working within enterprise-grade Cloud environments (AWS, Azure, or GCP) Distributed Querying: Experience with Trino (Presto SQL) Infrastructure as Code (IaC): Experience with Terraform Containerization: Experience with containers and container orchestration (Docker, Kubernetes) Managed Streaming: Practical experience with Confluent Cloud Scripting & Big Data Development: Practical experience with Python and Py.
Spark Scale: Prior experience working within a large Enterprise organization
Description copied from EPAM Systems's careers page. Read the full posting before you apply.
More jobs at EPAM Systems
Senior Full-Stack Engineer (Python+React)
EPAM Systems· Remote (Serbia)Manager/Senior Manager, Delivery Management - Application Security
EPAM Systems· London, England, UKSenior Data Scientist
EPAM Systems· Remote (Argentina; Brazil; Chile; Colombia; Mexico)Senior End-User Support Engineer
EPAM Systems· Warsaw, Masovian Voivodeship, PolandSenior Cloud & Infrastructure Engineer (Azure & Terraform)
EPAM Systems· Remote (Canada)