1,646 open roles
Databricks
Job description
Join a collaborative data engineering team where your work directly powers analytics, reporting, and data-driven decisions. In this role, you’ll help design and build reliable data pipelines on Databricks, transforming raw data into trusted, high-quality datasets that teams can confidently use. You’ll partner closely with stakeholders to understand data needs, improve performance, and ensure smooth end-to-end delivery—from ingestion to transformation and validation. If you enjoy solving real-world data challenges, optimizing distributed processing, and continuously improving how data platforms operate, this opportunity will help you grow your technical depth while contributing to impactful outcomes. You’ll work in an environment that values ownership, learning, and practical innovation—where clean engineering and teamwork go hand in hand.
Responsibilities
Key Responsibilities
• Develop and maintain scalable ETL pipelines using Databricks and PySpark to process large datasets efficiently. • Implement data transformations, cleansing, and enrichment logic aligned to business and analytics requirements. • Optimize Spark jobs for performance and cost by tuning partitions, caching, and cluster configurations where applicable. • Build reusable notebooks/jobs and support scheduling/orchestration of workloads within the Databricks environment. • Perform data validation, reconciliation, and quality checks to ensure accuracy and reliability of curated datasets. • Troubleshoot pipeline failures, analyze logs, and resolve issues to maintain stable production operations. • Collaborate with cross-functional teams to gather requirements, provide estimates, and deliver enhancements iteratively. • Maintain clear technical documentation for pipelines, transformations, and operational runbooks.
Requirements
• Primary skills:Technology->Data Engineering->Databricks
Additional:
Minimum Qualifications
• Bachelor’s degree (or equivalent) in Engineering/Technology/Computer Science or related field (BTech/BE/MSc or equivalent). • 3–5 years of experience in data engineering or related roles with hands-on Databricks experience. • Strong hands-on development experience with PySpark for distributed data processing. • Proven experience building and supporting ETL pipelines in production environments. • Ability to analyze data issues, debug Spark/ETL jobs, and implement reliable fixes.
Preferred Qualifications
• Experience designing end-to-end data workflows on Databricks including job scheduling, monitoring, and operational support. • Strong understanding of data modeling concepts and building curated datasets for analytics use cases. • Familiarity with Delta Lake concepts (ACID tables, incremental processing, upserts/merges) and best practices for lakehouse implementations. • Experience with performance tuning techniques for Spark workloads and handling large-scale datasets efficiently. • Exposure to CI/CD practices for data pipelines and maintaining code quality through reviews and standards. Good to have skills: Delta Lake, Spark SQL, Data Modeling, Workflow Orchestration, Performance Tuning
Description copied from Infosys's careers page. Read the full posting before you apply.
More jobs at Infosys
Aveva E3D / Aveva P&ID / AutoPIPE Engineer
Infosys· Noida, IndiaTechnical Publication Engineer
Infosys· Bangalore, IndiaDesign Engineer
Infosys· Bangalore, IndiaDesign Engineer (Medical Devices)
Infosys· Pune, IndiaCAD Engineer (Medical Devices)
Infosys· Pune, India
More jobs in Bengaluru
Developer (ACL)
L&T Technology Services· Bangalore, INFront End Developer
Machstatz· Bangalore, IndiaOryxion Screening
Oryxion· Bangalore, IndiaCustomer Experience Champion
MiStay· Bangalore, IndiaProduct Marketing Manager II
Docket Inc· Bengaluru, Karnataka, IN