Senior Associate Data Engineering
National Payments Corporation of India
Job description
The Opportunity To design and implement real-time data streaming pipelines ensuring 24×7 availability of high-volume transactional data.
Your role
includes building fault-tolerant, scalable systems using modern big data technologies, integrating data from multiple sources into data lakes/lakehouses, and enabling downstream analytics. You will collaborate with internal teams (Data Science, Product, and Technology) to deliver reliable data solutions, optimize performance, and maintain compliance with NPCI standards.
Job Details Job Title: Data Engineer (Streaming) Division: Data Analytics Years of Experience: 3–8 years Education: Graduation in Computer Science/IT (preferably BE/B.Tech) or equivalent; advanced degrees are a plus Employment Type: Full-time, Permanent Location: Mumbai & Hyderabad Key Responsibilities Design and develop real-time data pipelines using Apache Kafka and stream processing frameworks (Spark Structured Streaming / Apache Flink).
Ensure 24×7 data availability with fault-tolerant, highly reliable systems. Implement ingestion, transformation, and load (ITL) patterns for data lakes/lakehouses (S3/MinIO, HDFS). Work with table formats like Iceberg, Hudi, or Paimon for ACID transactions and schema evolution. Optimize SQL queries on Trino/Hive for large-scale analytics.
Develop orchestration workflows using DBT, Dagster, or Airflow for data transformations. Write efficient code in Python, Scala, and Java for data processing and automation. Collaborate with cross-functional teams to understand requirements and deliver high-quality data solutions. Monitor pipeline health, manage checkpoints, and implement observability for streaming jobs.
Ensure compliance with security, governance, and audit standards.
Requirements
Key Skills and Experience Required Mandatory Technical Skills: Apache Kafka (topics, partitions, offsets, reliability) Stream processing (Spark Structured Streaming or Apache Flink) Preferred Technical Skills: Data Lake / Lakehouse (S3, MinIO, HDFS) Table formats: Iceberg, Hudi, Paimon SQL engines: Trino, Hive Orchestration tools: DBT, Dagster, Airflow Programming: Python, Scala, Java NoSQL databases (MongoDB, Cassandra, Redis) CI/CD (Jenkins, GitHub Actions), Linux basics BI tools: Superset, Tableau Data Quality frameworks and observability practices Other Requirements: Strong analytical and problem-solving skills Ability to work in 24×7 environments Quick adaptability to new technologies