1,646 open roles
Hadoop / PySpark
Job description
Step into a high-impact consulting role where you’ll lead the design and delivery of scalable big data solutions that turn complex datasets into actionable insights. As a Lead Consultant, you’ll collaborate closely with data engineers, architects, analysts, and business stakeholders to shape modern data platforms built on Hadoop ecosystems and PySpark-based processing. You’ll guide teams through best practices, performance tuning, and reliable delivery—balancing hands-on problem solving with technical leadership. This role is ideal for someone who enjoys mentoring, driving technical decisions, and building data pipelines that are resilient, efficient, and production-ready. If you’re motivated by solving large-scale data challenges and enabling teams to deliver measurable outcomes, you’ll find a collaborative environment here that values ownership, clarity, and continuous improvement.
Responsibilities
• Lead end-to-end delivery of big data solutions using Hadoop and PySpark, from requirements to production rollout.
• Design and implement scalable batch/ETL pipelines for large datasets with strong focus on reliability and performance.
• Drive technical architecture discussions, define standards, and ensure best practices for distributed data processing.
• Optimize Spark jobs through partitioning strategies, caching, shuffle tuning, and efficient file formats where applicable.
• Collaborate with stakeholders to translate business needs into technical specifications, delivery plans, and milestones.
• Establish data quality checks, validation frameworks, and operational monitoring for production pipelines.
• Perform root-cause analysis for pipeline failures and performance bottlenecks; implement preventive fixes.
• Mentor engineers through code reviews, design reviews, and knowledge sharing to uplift team capability.
• Ensure documentation, runbooks, and handover artifacts are created and maintained for support readiness.
Requirements
• Primary skills:Technology->Big Data - Data Processing->PySpark,Technology->Big Data - Hadoop->Hadoop
Additional: • Bachelor’s or Master’s degree (or equivalent) in Engineering/Technology/Computer Applications/Science (BE/BTech/MTech/MCA/MSc or equivalent).
• 9–11 years of overall experience in data engineering and large-scale data processing environments.
• Strong hands-on experience with Hadoop ecosystem components and distributed data processing concepts.
• Strong hands-on experience building data pipelines using PySpark.
• Proven ability to lead technical delivery, guide teams, and manage stakeholder expectations in consulting engagements.
Preferred Qualifications
• Experience designing and implementing Spark-based data processing patterns (batch and incremental loads).
• Strong understanding of data modeling and storage patterns for big data platforms (partitioning, compaction, schema evolution).
• Experience with workflow orchestration and scheduling for data pipelines and dependency management.
• Demonstrated expertise in production hardening: monitoring, alerting, SLAs, and incident management for data jobs.
• Strong consulting mindset with ability to present solutions, document designs clearly, and influence technical decisions across teams.
Description copied from Infosys's careers page. Read the full posting before you apply.
More jobs at Infosys
Aveva E3D / Aveva P&ID / AutoPIPE Engineer
Infosys· Noida, IndiaTechnical Publication Engineer
Infosys· Bangalore, IndiaDesign Engineer
Infosys· Bangalore, IndiaDesign Engineer (Medical Devices)
Infosys· Pune, IndiaCAD Engineer (Medical Devices)
Infosys· Pune, India
More jobs in Bengaluru
Staff Software Engineer
Okta· Bengaluru, IndiaIntern - User Experience
ZF Group· Bangalore, KA, IN, 560058Senior Tax Analyst - Indirect Tax (Sales & Use Tax)
Expedia Group· India - BangaloreTech Lead - Enterprise Application ServiceNow N 4C
Genpact· 1401-GIPL: Prestige Technology Park IV, Bangalore