kdataai

Lead Databricks ML & AI Ops Engineer

kdataai

ContractPosted Jul 3, 2026

Job description

We are seeking a Lead Databricks ML & AI Ops Engineer to architect, scale, and govern the machine learning and AI platforms that power our enterprise data products. In this role, you will define the technical vision and engineering strategy for our end-to-end MLOps lifecycle—spanning from exploratory data science to global-scale production deployment on Databricks, Delta Lake, and Unity Catalog.

As a Lead Engineer, you will serve as the primary architect and technical authority for our ML platform. You will lead and mentor a talented team of senior engineers , collaborate closely with data science executives, and align with product teams to deliver robust, high-performance AI features. This role requires a balance of advanced distributed computing expertise, cutting-edge Generative AI platform design, and a strong DevOps/GitOps mindset.

Key Responsibilities

Technical Leadership & Strategy Platform Vision: Define the architectural roadmap for the enterprise ML platform, steering the migration away from legacy systems (e.g., Apache Airflow) to modern Databricks Workflows and Asset Bundles. Standards & Governance: Establish, document, and enforce global engineering standards for code quality, CI/CD pipelines, and automated testing.

Lead enterprise-wide ML governance and data security strategies utilizing Unity Catalog. Team Mentorship: Lead, mentor, and coach senior and mid-level engineers. Conduct advanced design and code reviews to foster a culture of technical excellence. Vendor & Community Engagement: Act as the primary technical point of contact for Databricks account teams.

Represent the company at external conferences and contribute to tech blogs. ML & LLM Platform Architecture Enterprise Pipelines: Architect reusable, highly efficient batch and streaming feature pipelines using Delta Live Tables, Spark Structured Streaming, and Databricks Feature Store. Compute Optimization: Standardize cluster configurations (including multi-node GPU training) to optimize complex AutoML and deep learning workloads (Py.

Torch, Hugging Face) for cost and performance. Generative AI Infrastructure: Design secure, production-grade frameworks for Large Language Model (LLM) operations (LLMOps), including robust RAG pipelines, agentic workflows, and Vector Search indexing using Mosaic AI and Lang. Chain. Model Lifecycle Management: Oversee the global setup of MLflow, defining enterprise staging gates, automated retraining loops, and instant rollback strategies.

Infrastructure & Operations (Ops) Production Serving: Architect high-availability, low-latency model serving frameworks utilizing Databricks Model Serving, Mosaic AI Gateway, and containerized deployments via Kubernetes/FastAPI. Infrastructure-as-Code: Lead the standardization of Infrastructure-as-Code (IaC) using Terraform or Pulumi to automate environment provisioning across multiple cloud regions.

Observability & SLAs: Establish advanced, proactive monitoring for data drift, model performance decay, hallucination rates, and pipeline SLAs using Databricks Lakehouse Monitoring, Prometheus, and Grafana.

Requirements

8+ years of software, data, or DevOps engineering experience, with at least 5 years dedicated specifically to production-grade ML/AI systems. Proven experience in a technical lead or architectural capacity. Databricks Mastery: Deep, expert-level knowledge of the Databricks ecosystem, including Workflows, Delta Live Tables, Unity Catalog, Mosaic AI, and Databricks Asset Bundles.

Distributed Systems: Advanced understanding of Spark internals (DAG optimization, shuffle tuning, memory management) at an enterprise, multi-terabyte scale. Core Engineering: Exceptional Python skills (writing highly optimized, type-annotated, and modular code) alongside deep familiarity with frameworks like Py. Torch, scikit-learn, and XGBoost.

DevOps & GitOps: Strong background in enterprise Git. Ops workflows (e.g., trunk-based development), Docker containerization, Kubernetes orchestration, and complex GitHub Actions CI/CD pipelines. Cloud Architecture: Extensive hands-on experience provisioning and securing cloud infrastructure on AWS, Azure, or GCP.

Preferred Qualifications

Databricks Certified Machine Learning Professional and/or Databricks Certified Enterprise Architect. Advanced AI/Governance: Direct experience implementing Responsible AI frameworks, model cards, bias auditing, and cost-tracking guardrails for LLMs. Open Source: Active contributor to open-source ML, MLOps, or data engineering projects (e.

g., MLflow, Delta Lake, Lang. Chain). Data Mesh: Experience designing Lakehouse architectures within a decentralized Data Mesh organizational framework. Technical Environment Platform & Storage: Databricks (AWS/Azure/GCP), Delta Lake, Unity Catalog, S3/ADLS Gen2. Orchestration & CI/CD: Databricks Workflows, Databricks Asset Bundles, GitHub Actions, Terraform.

ML & LLM Frameworks: MLflow, Py. Torch, Hugging Face, Mosaic AI, Lang. Chain, Databricks Vector Search, GPT-4 / Claude APIs. Serving & Observability: Databricks Model Serving, FastAPI, Docker, Kubernetes, Databricks Lakehouse Monitoring, Prometheus, Grafana.