China Minmetals

Senior Data Engineer

China Minmetals

Beijing, ChinaPosted Apr 8, 2026

Job description

Role Purpose: We are looking for an experienced Data Engineer with full-stack development capabilities to join our enterprise digital team. In this role, you will be responsible for designing, implementing, maintaining, and optimizing data pipelines and data products primarily within the Azure cloud ecosystem, with an emphasis on Azure Databricks and Azure Data Factory.

You will closely collaborate with SAP engineers, data analysts, other data engineers, business analysts, and stakeholders to deliver high-quality, efficient, and scalable data solutions supporting analytics, reporting, application development, and data science initiatives. As a Data Engineer, you will play a critical role in creating and managing robust, production-grade data pipelines that integrate multiple data sources into the Databricks Lakehouse.

Your responsibilities

will include maintaining and optimizing existing ELT frameworks and data pipelines, as well as designing and building new data pipelines and data models based on evolving business requirements. You will work closely with stakeholders to define and implement business metrics, and develop data visualization and reporting solutions using Power BI to support data-driven decision making.

In addition, you will ensure all data solutions adhere to established data governance, data quality, and security standards. At our company, data engineering is not just about moving data, but about building trusted and reusable data assets. We focus on delivering actionable data quickly, while proactively reducing data debt through strong governance, compliance, cataloging, and operational best practices.

Required Capabilities: Develop, manage, monitor and optimize scalable data pipelines (ELT processes) with a strong focus on data ingestion, reliability, performance, and scalability, using Azure Data Factory (ADF) and Databricks Lakeflow (Jobs and Pipelines) as the primary pipeline development and orchestration platforms within the Databricks Lakehouse.

Perform day-to-day operations, monitoring, and continuous performance optimization across the end-to-end data platform, including key components such as Azure Data Factory (ADF, including Self-hosted Integration Runtime), Azure Virtual Networks (VNET), Databricks, Azure Functions, Azure Storage Accounts, Azure Key Vault, Azure DevOps, and Power BI (including On-premises/VNET Data Gateway).

Ensure platform stability, reliability, security, and scalability to consistently support business-critical data workloads. Define, manage, and govern master data, transactional data, logical data models, metadata, and business metrics through a layered Delta Lake architecture. Leverage Unity Catalog (Databricks) to enforce data standards, fine-grained access control, auditing, and controlled data distribution, ensuring secure, compliant, and consistent data consumption across domains.

Apply best practices in scripting and automation, with strong proficiency in Python, Scala, and SQL for data processing and data product development. Familiarity with additional languages such as VBA or DEX (Power BI) considered a plus. Perform database administration and performance tuning for SQL-based databases, NoSQL databases, and Databricks Lakehouse solutions, with hands-on experience in optimizing workloads.

Collaborate with domain experts, data scientists, and analysts to translate business requirements into well-defined data models, pipelines, and metrics that support analytical and reporting needs. Troubleshoot and optimize data pipelines to maintain high performance, reliability, and scalability. Ensure rigorous adherence to data security, governance, and compliance guidelines.

Improve and maintain CI/CD pipelines and manage code repositories, particularly for our existing Python-based ELT framework in Azure DevOps. Prepare comprehensive technical documentation and maintain best practices across projects. Manage and minimize data debt, optimizing existing data platform functionalities. Bachelor’s or Master’s degree in Computer Science, Mathematics, Data Engineering, Information Systems, or a related technical field.

5-8 years of professional experience as a Data Engineer or in a similar capacity. 5-8 years of hands-on experience with SQL, Python, Spark SQL, and/or Py. Spark. Proven expertise (5-8 years) in cloud data platforms, preferably Azure Databricks and Azure Data Factory. Proficiency with big data frameworks and tools such as Apache Hadoop and Apache Spark.

5-8 years of experience working with SDLC project management tools like Azure DevOps. 3-5 years of experience with version control systems, especially Git. Experience with Infrastructure on Azure, better have hands-on on Code (IaC) tools, such as Terraform. Strong experience (3-5 years) using business intelligence tools such as Power BI.

Solid experience with ETL tools (Azure Data Factory) and processes. Familiarity with CI/CD pipelines and managing code repositories, particularly in Azure DevOps. Strong knowledge of data modelling, data warehousing, and data lake concepts. Demonstrated ability to handle large, complex datasets, including data harmonization to support detailed analytics, dashboards, and reporting.

Excellent analytical and problem-solving skills with high attention to detail. Strong communication skills with the ability to effectively collaborate with technical and non-technical stakeholders to gather requirements and deliver successful solutions. Experience and familiarity with Data Lakehouse/Delta Lake (Databricks).

Experience with data governance platform of Databricks. Experience with SAP data management and extraction. Experience with machine learning models. Experience with Azure cloud operations. Certification in Databricks or Azure. Experience working with Large Language Models (LLMs). Strong proficiency in English, with the ability to clearly understand, communicate, and effectively express technical and business-related concepts.

Relevant Azure and/or Databricks certifications are preferred.