Job description
Responsibilities
5+ years of experience utilizing Python, Py. Spark/Scala for processing large data volumes Good knowledge of working with Databases (e.g., SQL, Oracle, Hive, Impala) Relevant experience with Pandas, Num. Py, Scikit-learn, and other relevant Python libraries Google Cloud Platform (Data. Proc, Air. Flow, Google Cloud Storage) Good knowledge of Unix/Linux environments Familiarity with Big Data and Hadoop Ecosystem: Spark (Spark SQL, Dataframes, Py.
Spark, HUE, parquet files) Awareness about DevOps practices like CI/CD Proficient communication and English language skills (written/verbal) Skills Py. Spark (Python, Py. Spark or Scala) Spark Data. Frames ETL concepts Database and SQL knowledge (Hive/Impala)